The overfitting problem: deflated Sharpe and what discipline looks like
SIG · Chapter 812 min readAsked at Two Sigma, AQR, Citadel, Point72
Assumes Execution: market impact, implementation shortfall and capacity.
After this lesson you should be able to
- Compute the Sharpe ratio that luck alone produces from trials.
- Estimate the minimum backtest length for a claimed Sharpe.
- Describe the process controls that actually work.
This is the central problem of the field. Financial data are short and noisy, computers are fast, and a sufficiently determined search will find something in any sample. Everything else in signal research is a defence against that fact.
Equation 8.1
What luck produces
The expected best Sharpe ratio among worthless strategies over years. Any reported Sharpe must clear this before it is evidence of anything.
- Every variant tried — parameter settings, universes, horizons, abandoned attempts.
- Grows slowly, so the fix is more data far more than fewer trials.
| Trials | 2 years | 5 years | 10 years | 20 years |
|---|---|---|---|---|
| 10 | 1.52 | 0.96 | 0.68 | 0.48 |
| 100 | 2.15 | 1.36 | 0.96 | 0.68 |
| 1,000 | 2.63 | 1.66 | 1.18 | 0.83 |
| 10,000 | 3.03 | 1.92 | 1.36 | 0.96 |
Proposition 8.3
Minimum backtest length
Invert the formula: to claim a Sharpe of after trials, you need roughly years. A Sharpe of after a thousand trials needs about fourteen years — which is more history than most signals have, and is the honest reason so many published results do not replicate.
Holds when
- The requirement falls with the square of the Sharpe, so a genuinely strong signal needs far less data.
- It rises only logarithmically in the trial count, so cutting trials from 1,000 to 100 helps surprisingly little.
- More independent data is the only strong lever, which is why breadth across assets matters so much.
Why fewer trials helps less than you think. The threshold grows as , so going from ten thousand trials to a hundred lowers the bar by only about thirty per cent — which is a genuinely uncomfortable result, because reducing the search is the intervention everyone reaches for first. What it means is that discipline about the *number* of tests is necessary and nowhere near sufficient. The effective levers are longer or wider data, which enters as , and a stronger prior from an economic mechanism, which changes the base rate rather than the threshold.
The rest of this lesson is in Premium
You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.