Why k-fold cross-validation is wrong on financial data
ML · Chapter 512 min readAsked at Two Sigma, Citadel, QuantCo, Point72
Assumes p-values, p-hacking and the multiple-testing problem.
After this lesson you should be able to
- Explain the leaks that ordinary k-fold creates on a time series.
- Describe purging and embargoing.
- Say why tree ensembles dominate tabular financial problems and where they fail.
Standard machine-learning practice assumes the rows are exchangeable. Financial rows are not: they are ordered, overlapping and serially correlated. Applying k-fold cross-validation to them produces validation scores that are systematically too good, and the model that wins the comparison is usually the one that leaks best.
| Leak | What happens | The fix |
|---|---|---|
| Temporal | Training on rows after the validation rows | Split by time; train only on the past |
| Overlap | A label spanning several bars appears on both sides of the split | Purge rows whose label window crosses the boundary |
| Residual correlation | Rows just after the boundary still carry the same information | Embargo a gap after the validation set |
Definition 5.2
Purging and embargoing
Purged, embargoed cross-validation — Purging drops training rows whose label window overlaps the validation window. Embargoing additionally drops a short block of training rows immediately *after* the validation set, because serial correlation means those rows still carry information about it. Together they make a cross-validated score on overlapping labels honest, at the cost of some training data.
Proposition 5.3
Walk-forward is the honest default
Train on a window, test on the block that follows, roll forward, repeat. It is closer to how the model will actually be used than any shuffled scheme, it surfaces regime dependence, and it makes performance decay over time visible rather than averaging it away.
Holds when
- Expanding window if you believe old data stays relevant; rolling window if you believe the regime shifts.
- Report the per-period results, not just the average — an average hides a strategy that worked only in 2009.
- Every hyperparameter chosen on validation data is a test, and it counts toward the trial budget.
The rest of this lesson is in Premium
You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.