Forecast evaluation: walk-forward design and backtest overfitting
TS · Chapter 713 min readAsked at Two Sigma, AQR, Citadel, Point72
Assumes Stylised facts: what financial returns actually look like.
After this lesson you should be able to
- Design a walk-forward evaluation that does not leak.
- Compare two forecasts with a Diebold–Mariano test.
- Deflate a Sharpe ratio for the number of trials behind it.
Evaluating a forecast on financial data is harder than producing one. The design has to respect time, the comparison has to account for the fact that both forecasts saw the same data, and the final number has to be discounted for how many candidates were tried to get it.
Proposition 7.2
Walk-forward design
Fit on a window, test on the block that follows, roll forward, repeat. Every parameter — including the hyperparameters and the feature selection — must be chosen using only data from before the test block, or the evaluation is contaminated.
Holds when
- Expanding window if the data-generating process is stable; rolling if you expect regime change.
- Report the per-period results, not just the average — a strategy that worked only in one regime should be visible.
- Anything fitted outside the loop, including a scaling factor or a universe filter, leaks.
| Leak | How it hides | Fix |
|---|---|---|
| Hyperparameters tuned on all data | Looks like one model, is really hundreds | Nest the tuning inside the loop |
| Full-sample normalisation | No code reads a future row | Expanding or cross-sectional statistics |
| Universe chosen today | Survivorship | Point-in-time membership |
| Overlapping labels | Train and test share information | Purge and embargo |
| Reused test set | Each look is a trial | Touch the final holdout once |
Equation 7.4
Diebold–Mariano
Test whether two forecasts differ in accuracy by testing whether their loss differential has mean zero.
- The loss function — squared error, absolute error, or something economic.
- Must be HAC, since the loss differential is autocorrelated.
Why you cannot just compare two error rates. Both forecasts were evaluated on the same periods, so their errors are heavily correlated — when the market gapped, both were wrong. Comparing their mean squared errors independently ignores that and produces a standard error far too large, so real differences look insignificant. Diebold–Mariano tests the *difference* series directly, where the common component has already cancelled, which is why it can detect improvements that a naive comparison misses entirely.
The rest of this lesson is in Premium
You have read the opening. 11 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.