Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
      • 1Framework

        • The framework: bias, variance, capacity and dimensionality
      • 2Tree-based methods

        • Trees: bagging, random forests and gradient boosting
      • 3Other supervised methods

        • Other supervised methods: kNN, SVMs and the kernel trick
      • 4Unsupervised learning

        • Unsupervised learning: clustering assets and correlation structure
      • 5Model selection and evaluation

        • Why k-fold cross-validation is wrong on financial data
      • 6Optimisation for learning

        • Optimisation: gradient descent, momentum and Adam
      • 7Neural networks

        • Neural networks: backpropagation, and when they are the wrong tool
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Machine learning
  4. /Model selection and evaluation

Why k-fold cross-validation is wrong on financial data

ML · Chapter 5·12 min read·Asked at Two Sigma, Citadel, QuantCo, Point72

Assumes p-values, p-hacking and the multiple-testing problem.

After this lesson you should be able to

  • Explain the leaks that ordinary k-fold creates on a time series.
  • Describe purging and embargoing.
  • Say why tree ensembles dominate tabular financial problems and where they fail.

Standard machine-learning practice assumes the rows are exchangeable. Financial rows are not: they are ordered, overlapping and serially correlated. Applying k-fold cross-validation to them produces validation scores that are systematically too good, and the model that wins the comparison is usually the one that leaks best.

LeakWhat happensThe fix
TemporalTraining on rows after the validation rowsSplit by time; train only on the past
OverlapA label spanning several bars appears on both sides of the splitPurge rows whose label window crosses the boundary
Residual correlationRows just after the boundary still carry the same informationEmbargo a gap after the validation set
Table 5.1 · The leaks. The second row is the one that catches experienced practitioners. A five-day forward return computed every day means each label overlaps the next four, so a naive split shares information across it even when the dates do not overlap.

Definition 5.2

Purging and embargoing

Purged, embargoed cross-validation — Purging drops training rows whose label window overlaps the validation window. Embargoing additionally drops a short block of training rows immediately *after* the validation set, because serial correlation means those rows still carry information about it. Together they make a cross-validated score on overlapping labels honest, at the cost of some training data.

Proposition 5.3

Walk-forward is the honest default

Train on a window, test on the block that follows, roll forward, repeat. It is closer to how the model will actually be used than any shuffled scheme, it surfaces regime dependence, and it makes performance decay over time visible rather than averaging it away.

Holds when

  • Expanding window if you believe old data stays relevant; rolling window if you believe the regime shifts.
  • Report the per-period results, not just the average — an average hides a strategy that worked only in 2009.
  • Every hyperparameter chosen on validation data is a test, and it counts toward the trial budget.
Usable for training900Purged at the fold boundaries100
Figure 5.4 · What purging costs you, on 1,000 days. Ten folds, a label that takes five days to resolve, and five days lost either side of every boundary. A tenth of the sample goes — and paying it is the difference between a cross-validated number that means something and one that has read the answer.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Unsupervised learning: clustering assets and correlation structureOptimisation: gradient descent, momentum and Adam →
On this page
  • The leaks
  • Purging and embargoing
  • Walk-forward is the honest default
  • What purging costs you, on 1,000 days

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.