Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
      • 1The research pipeline

        • The research pipeline: from hypothesis to live capital
      • 2Signal construction

        • Constructing a signal: standardisation, neutralisation and combination
      • 3Measuring a signal

        • Measuring a signal: IC, breadth and the fundamental law
      • 4Factor models

        • Factor models: CAPM, Fama–French and statistical factors
      • 5Risk models

        • Risk models: covariance estimation, VaR and expected shortfall
      • 6Portfolio construction

        • Portfolio construction: mean-variance, and why nobody uses it raw
      • 7Execution and costs

        • Execution: market impact, implementation shortfall and capacity
      • 8The overfitting problem

        • The overfitting problem: deflated Sharpe and what discipline looks like
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Alpha and signal research
  4. /The overfitting problem

The overfitting problem: deflated Sharpe and what discipline looks like

SIG · Chapter 8·12 min read·Asked at Two Sigma, AQR, Citadel, Point72

Assumes Execution: market impact, implementation shortfall and capacity.

After this lesson you should be able to

  • Compute the Sharpe ratio that luck alone produces from NNN trials.
  • Estimate the minimum backtest length for a claimed Sharpe.
  • Describe the process controls that actually work.

This is the central problem of the field. Financial data are short and noisy, computers are fast, and a sufficiently determined search will find something in any sample. Everything else in signal research is a defence against that fact.

Equation 8.1

What luck produces

The expected best Sharpe ratio among NNN worthless strategies over TTT years. Any reported Sharpe must clear this before it is evidence of anything.

E[max⁡i≤NSRi]≈2ln⁡NT\mathbb{E}\left[\max_{i \le N} \mathrm{SR}_i\right] \approx \sqrt{\frac{2\ln N}{T}}E[i≤Nmax​SRi​]≈T2lnN​​
NNN
Every variant tried — parameter settings, universes, horizons, abandoned attempts.
2ln⁡N\sqrt{2\ln N}2lnN​
Grows slowly, so the fix is more data far more than fewer trials.
Trials2 years5 years10 years20 years
101.520.960.680.48
1002.151.360.960.68
1,0002.631.661.180.83
10,0003.031.921.360.96
Table 8.2 · The luck threshold. Read the row for how many things you tried and the column for how much data you have. A Sharpe of 1.5 on two years is below what luck gives after a hundred attempts, and comfortably above the bar after ten attempts on twenty years.

Proposition 8.3

Minimum backtest length

Invert the formula: to claim a Sharpe of SSS after NNN trials, you need roughly T≥2ln⁡N/S2T \ge 2\ln N / S^2T≥2lnN/S2 years. A Sharpe of 111 after a thousand trials needs about fourteen years — which is more history than most signals have, and is the honest reason so many published results do not replicate.

Holds when

  • The requirement falls with the square of the Sharpe, so a genuinely strong signal needs far less data.
  • It rises only logarithmically in the trial count, so cutting trials from 1,000 to 100 helps surprisingly little.
  • More independent data is the only strong lever, which is why breadth across assets matters so much.
125005000024√(2 ln N)Backtests runExpected best Sharpe, all of them worthless
Figure 8.4 · The best of nothing. With a thousand strategies that have no edge at all, the best one shows a Sharpe near 3.73.73.7 by luck alone. This is why the number of configurations tried is part of the result: without it, a Sharpe is not evidence of anything.

Why fewer trials helps less than you think. The threshold grows as ln⁡N\sqrt{\ln N}lnN​, so going from ten thousand trials to a hundred lowers the bar by only about thirty per cent — which is a genuinely uncomfortable result, because reducing the search is the intervention everyone reaches for first. What it means is that discipline about the *number* of tests is necessary and nowhere near sufficient. The effective levers are longer or wider data, which enters as T\sqrt{T}T​, and a stronger prior from an economic mechanism, which changes the base rate rather than the threshold.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Execution: market impact, implementation shortfall and capacityBack to Alpha and signal research →
On this page
  • What luck produces
  • The luck threshold
  • Minimum backtest length
  • The best of nothing

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.