Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
    • TSTime series
      • 1Foundations

        • Foundations: stationarity, autocorrelation and the Wold decomposition
      • 2The ARMA family

        • The ARMA family: identification, estimation and forecasting
      • 3Unit roots and cointegration

        • Unit roots, spurious regression and the basis of pairs trading
      • 4Volatility modelling

        • Volatility models: ARCH, GARCH and realised measures
      • 5State space and filtering

        • State space and the Kalman filter
      • 6Financial stylised facts

        • Stylised facts: what financial returns actually look like
      • 7Forecast evaluation

        • Forecast evaluation: walk-forward design and backtest overfitting
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Time series
  4. /Forecast evaluation

Forecast evaluation: walk-forward design and backtest overfitting

TS · Chapter 7·13 min read·Asked at Two Sigma, AQR, Citadel, Point72

Assumes Stylised facts: what financial returns actually look like.

After this lesson you should be able to

  • Design a walk-forward evaluation that does not leak.
  • Compare two forecasts with a Diebold–Mariano test.
  • Deflate a Sharpe ratio for the number of trials behind it.

Evaluating a forecast on financial data is harder than producing one. The design has to respect time, the comparison has to account for the fact that both forecasts saw the same data, and the final number has to be discounted for how many candidates were tried to get it.

Training500Purge5Embargo5Test100
Figure 7.1 · One walk-forward window, in days. The two small blocks are the whole point. A label that takes five days to resolve reaches five days into the future, so the five days either side of the boundary have to be thrown away — and a backtest that keeps them is reporting a number it did not earn.

Proposition 7.2

Walk-forward design

Fit on a window, test on the block that follows, roll forward, repeat. Every parameter — including the hyperparameters and the feature selection — must be chosen using only data from before the test block, or the evaluation is contaminated.

Holds when

  • Expanding window if the data-generating process is stable; rolling if you expect regime change.
  • Report the per-period results, not just the average — a strategy that worked only in one regime should be visible.
  • Anything fitted outside the loop, including a scaling factor or a universe filter, leaks.
LeakHow it hidesFix
Hyperparameters tuned on all dataLooks like one model, is really hundredsNest the tuning inside the loop
Full-sample normalisationNo code reads a future rowExpanding or cross-sectional statistics
Universe chosen todaySurvivorshipPoint-in-time membership
Overlapping labelsTrain and test share informationPurge and embargo
Reused test setEach look is a trialTouch the final holdout once
Table 7.3 · Where evaluations leak. The first and the last are the same failure at different scales: every decision informed by the test data is a use of it, whether or not any code reads a future row.

Equation 7.4

Diebold–Mariano

Test whether two forecasts differ in accuracy by testing whether their loss differential has mean zero.

DM=dˉVar^(dˉ),dt=L(e1t)−L(e2t)DM = \frac{\bar{d}}{\sqrt{\widehat{\mathrm{Var}}(\bar{d})}}, \qquad d_t = L(e_{1t}) - L(e_{2t})DM=Var(dˉ)​dˉ​,dt​=L(e1t​)−L(e2t​)
LLL
The loss function — squared error, absolute error, or something economic.
Var^\widehat{\mathrm{Var}}Var
Must be HAC, since the loss differential is autocorrelated.

Why you cannot just compare two error rates. Both forecasts were evaluated on the same periods, so their errors are heavily correlated — when the market gapped, both were wrong. Comparing their mean squared errors independently ignores that and produces a standard error far too large, so real differences look insignificant. Diebold–Mariano tests the *difference* series directly, where the common component has already cancelled, which is why it can detect improvements that a naive comparison misses entirely.

The rest of this lesson is in Premium

You have read the opening. 11 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Stylised facts: what financial returns actually look likeBack to Time series →
On this page
  • One walk-forward window, in days
  • Walk-forward design
  • Where evaluations leak
  • Diebold–Mariano

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.