Skip to content
QuantMax
QuantMax
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
      • 1Limit theorems

        • The law of large numbers and the central limit theorem
      • 2Estimation

        • Estimation: bias, variance and maximum likelihood
      • 3Confidence intervals

        • Confidence intervals, and the interval you would trade
      • 4Hypothesis testing

        • Hypothesis testing: errors, power and which test to use
      • 5p-values and multiple testing

        • p-values, p-hacking and the multiple-testing problem
      • 6Resampling

        • Resampling: the bootstrap, permutation tests and where they break
      • 7Bayesian statistics

        • Bayesian statistics: conjugacy, shrinkage and credible intervals
      • 8Experiment design

        • Experiment design: randomisation, peeking and minimum detectable effect
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Statistics and inference
  4. /Limit theorems

The law of large numbers and the central limit theorem

STAT · Chapter 1·12 min read·Asked at Two Sigma, Citadel, DE Shaw, Jane Street

After this lesson you should be able to

  • State what each theorem says, and what it does not.
  • Use the n\sqrt{n}n​ rate to size a sample or a backtest.
  • Name the conditions under which both fail on financial data.

Two theorems carry most of applied statistics. The law of large numbers says an average converges to the truth; the central limit theorem says how fast and in what shape. Almost every question about sample size, error bars or the significance of a backtest is one of these in disguise.

Definition 1.1

The law of large numbers

LLN, Xˉn→μas n→∞\bar{X}_n \to \mu \quad \text{as } n \to \inftyXˉn​→μas n→∞ — The sample mean of independent draws with a finite mean converges to that mean. The weak form says it converges in probability, the strong form almost surely; for interview purposes the distinction rarely matters, but knowing there is one does.

Equation 1.2

The central limit theorem

The standardised sample mean converges in distribution to a standard normal, whatever the shape of the underlying distribution — provided the variance is finite.

n Xˉn−μσ →d N(0,1)\sqrt{n}\,\frac{\bar{X}_n - \mu}{\sigma} \ \xrightarrow{d}\ \mathcal{N}(0, 1)n​σXˉn​−μ​ d​ N(0,1)
σ/n\sigma/\sqrt{n}σ/n​
The standard error: the standard deviation of the sample mean.
→d\xrightarrow{d}d​
Convergence in distribution — the shape converges, not the individual values.

Proposition 1.3

Everything costs four times as much

Error shrinks as 1/n1/\sqrt{n}1/n​, so halving your error bar takes four times the data. This single fact sets the economics of research: an effect you can measure to one decimal place in a year of data needs a century for two. It is also why Monte Carlo is slow and why short backtests prove very little.

Holds when

  • To cut a confidence interval in half, quadruple the sample.
  • Monte Carlo error is O(1/n)O(1/\sqrt{n})O(1/n​) regardless of dimension, which is why it beats quadrature in high dimensions and loses in low ones.
110020030040000.511/√n1/nSample size nStandard error
Figure 1.4 · Why the error falls like the square root. Four times the data halves the error, not quarters it. That is the whole cost structure of estimation: the dashed line is what people assume they are getting, and the solid one is what averaging actually buys.

Example 1.5

A strategy has a true Sharpe ratio of 111. Roughly how many years of daily returns do you need before the estimated Sharpe is significant at two standard errors?

Show the worked solutionHide the worked solution

Worked solution

  1. Formula
    se(SR^)≈1T(T in years)\mathrm{se}(\widehat{SR}) \approx \frac{1}{\sqrt{T}} \quad (T \text{ in years})se(SR)≈T​1​(T in years)
  2. Substitute
    Need SRse=SRT≥2\text{Need } \frac{SR}{\mathrm{se}} = SR\sqrt{T} \ge 2Need seSR​=SRT​≥2
  3. Solve
    1×T≥21 \times \sqrt{T} \ge 21×T​≥2
  4. T≥4T \ge 4T≥4
  5. Answer
    about 4 years\text{about 4 years}about 4 years

Sanity check. A Sharpe of 0.5 would need sixteen years, which is the honest reason weak signals are so hard to validate.

ViolationWhat happensWhere it shows up
Infinite varianceCLT does not apply; the average does not concentrateCauchy-like tails, ratios of returns
Infinite meanEven the LLN failsThe St. Petersburg payoff; pure Pareto with α≤1\alpha \le 1α≤1
DependenceEffective sample size is far below nnnAutocorrelated returns, overlapping windows
Non-identical distributionsConvergence may still hold, but slowlyRegime change, structural breaks
Table 1.6 · When the theorems fail. The third row is the one that bites in practice: overlapping 12-month returns sampled monthly look like hundreds of observations and behave like a few dozen.

Why normality appears at all. A sum forgets the details of its parts. Adding many independent contributions, each small relative to the total, washes out everything except the mean and the variance — the higher moments of the individual terms scale away faster than the sum grows. That is why the normal turns up everywhere, and equally why it stops turning up when a single term can dominate the sum, which is exactly what a fat tail is.

Common trap. Claiming the CLT says the data become normal. It says nothing about the data; it is a statement about the sampling distribution of an average. Daily equity returns are not normal and never become so no matter how many you collect. Instead. Say what converges: the standardised *mean*, not the observations. Averages of returns are much closer to normal than the returns themselves.

Example 1.7

How fast the CLT arrives

Daily returns are heavily skewed. You average 202020 of them and then 250250250. How much does the skew of the mean fall?

Show the worked solutionHide the worked solution

Worked solution

  1. Formula
    skew(Xˉn)=skew(X)n\text{skew}(\bar X_n) = \frac{\text{skew}(X)}{\sqrt{n}}skew(Xˉn​)=n​skew(X)​
  2. Substitute
    n=20→250n = 20 \to 250n=20→250
  3. Solve
    20=4.47, 250=15.81\sqrt{20} = 4.47, \ \sqrt{250} = 15.8120​=4.47, 250​=15.81
  4. skew falls by 4.47× then 15.81×\text{skew falls by } 4.47\times \text{ then } 15.81\timesskew falls by 4.47× then 15.81×
  5. Answer
    A year is 3.5× closer to normal than a month\text{A year is } 3.5\times \text{ closer to normal than a month}A year is 3.5× closer to normal than a month

Sanity check. Skew dies like 1/n1/\sqrt{n}1/n​ and kurtosis like 1/n1/n1/n, which is why monthly return series look normal and daily ones never do — and why a test calibrated on monthly data misfires on daily.

Example 1.8

The two theorems answer different questions

You toss a fair coin a million times. How close is the proportion to a half, and how many heads should you expect the count to be away from 500,000500{,}000500,000?

Show the worked solutionHide the worked solution

Worked solution

  1. Formula
    sd(p^)=p(1−p)n,sd(count)=np(1−p)\text{sd}(\hat p) = \sqrt{\tfrac{p(1-p)}{n}}, \quad \text{sd}(\text{count}) = \sqrt{np(1-p)}sd(p^​)=np(1−p)​​,sd(count)=np(1−p)​
  2. Substitute
    n=106n = 10^6n=106
  3. Solve
    sd(p^)=0.51000=0.0005\text{sd}(\hat p) = \tfrac{0.5}{1000} = 0.0005sd(p^​)=10000.5​=0.0005
  4. sd(count)=500\text{sd}(\text{count}) = 500sd(count)=500
  5. Answer
    The proportion converges; the count diverges\text{The proportion converges; the count diverges}The proportion converges; the count diverges

Sanity check. The fraction gets closer to a half and the absolute number of surplus heads gets further from zero. Both are true and candidates who have only heard of the law of large numbers get the second one wrong.

Equation 1.9

Chebyshev’s inequality

Holds for any distribution with a finite variance — no normality required. It is loose, but it is the guarantee you fall back on when you cannot trust the tails, and it is what proves the weak law of large numbers in one line.

Pr⁡(∣X−μ∣≥kσ)≤1k2\Pr\big(|X - \mu| \ge k\sigma\big) \le \frac{1}{k^2}Pr(∣X−μ∣≥kσ)≤k21​

Example 1.10

A guarantee without normality

A sample mean has standard error σ/n\sigma/\sqrt nσ/n​. Without assuming normality, what is the most you can say about the chance it lands two standard errors or more from the truth?

Show the worked solutionHide the worked solution

Worked solution

  1. Formula
    Pr⁡(∣Xˉ−μ∣≥2 SE)≤122\Pr(|\bar X - \mu| \ge 2\,\mathrm{SE}) \le \frac{1}{2^2}Pr(∣Xˉ−μ∣≥2SE)≤221​
  2. Substitute
    k=2k = 2k=2
  3. Solve
    ≤0.25\le 0.25≤0.25
  4. Answer
    at most 25%\text{at most } 25\%at most 25%

Sanity check. Under normality the answer is 4.6%4.6\%4.6%. The gap between the two is the price of making no assumption about the tails — and heavy-tailed returns sit somewhere between.

Equation 1.11

The standard error of a Sharpe ratio

For independent, normal returns, measured per period (Lo, 2002). Annualise both the ratio and its error by periods per year\sqrt{\text{periods per year}}periods per year​. Fat tails and autocorrelation both make the true error larger.

SE⁡(SR^)≈1+12SR2n\operatorname{SE}(\widehat{SR}) \approx \sqrt{\frac{1 + \tfrac{1}{2} SR^2}{n}}SE(SR)≈n1+21​SR2​​

Example 1.12

How precise is four years of Sharpe?

A strategy’s daily Sharpe ratio is 0.10.10.1 over 1,0001{,}0001,000 days. What is the standard error of its annualised Sharpe ratio?

Show the worked solutionHide the worked solution

Worked solution

  1. Formula
    SE⁡daily=1+12(0.1)21000\operatorname{SE}_{\text{daily}} = \sqrt{\frac{1 + \tfrac{1}{2}(0.1)^2}{1000}}SEdaily​=10001+21​(0.1)2​​
  2. Substitute
    1.0051000=0.0317\sqrt{\frac{1.005}{1000}} = 0.031710001.005​​=0.0317
  3. Solve
    annualise×252: 0.0317×15.87\text{annualise} \times \sqrt{252}:\ 0.0317 \times 15.87annualise×252​: 0.0317×15.87
  4. Answer
    ≈0.50 on an annual Sharpe of 1.59\approx 0.50 \text{ on an annual Sharpe of } 1.59≈0.50 on an annual Sharpe of 1.59

Sanity check. A two-standard-error interval runs from about 0.60.60.6 to 2.62.62.6 — four years of data pin a good strategy’s Sharpe down only loosely.

Proposition 1.13

When the CLT does not apply

The central limit theorem needs a finite variance. For a Student-ttt with three degrees of freedom the variance is finite and averages do become normal, but slowly; with two or fewer the variance is infinite, and suitably scaled sums converge to a stable, non-normal law instead. Returns with tail exponents near three are why "approximately normal averages" can take far more data than the textbook suggests.

Holds when

  • Finite variance: CLT, at rate 1/n1/\sqrt n1/n​, slower with heavy tails and skew.
  • Infinite variance: no CLT; sums scale like n1/αn^{1/\alpha}n1/α with α<2\alpha < 2α<2.

How fast normality arrives. The Berry–Esseen theorem bounds the gap between the distribution of a standardised mean and the normal by a constant times the skewness over n\sqrt nn​. Symmetric data become normal quickly; skewed data — option payoffs, default losses — can need thousands of observations before tail probabilities from a normal approximation can be trusted.

Common trap — averaging ratios. Averaging per-period Sharpe ratios, or averaging returns on different capital bases, and calling it the ratio for the whole. The average of ratios is not the ratio of averages. Instead. Compute the statistic on the pooled data — total mean over total standard deviation — or weight the pieces properly. For returns, average log returns or compound the simple ones.

What you need to know

  • LLN: the sample mean converges to the true mean, given a finite mean.
  • CLT: the standardised sample mean converges to a standard normal, given a finite variance.
  • The standard error is σ/n\sigma/\sqrt{n}σ/n​ — quartering the error takes sixteen times the data.
  • The CLT is about the distribution of an average, not about the data.
  • Dependence is the practical killer: it shrinks the effective sample size without shrinking nnn.

Exercise 1.14

You compute a monthly Sharpe ratio from overlapping twelve-month returns sampled every month over ten years. Why is the resulting error bar too small?

Show the answerHide the answer

Consecutive observations share eleven of their twelve months, so they are heavily correlated. You have 120 numbers but roughly ten independent ones, and dividing by 120\sqrt{120}120​ rather than 10\sqrt{10}10​ understates the standard error by a factor of about three and a half.

Exercise 1.15

Three standard deviations

What does Chebyshev guarantee about a move of at least three standard deviations, and what would a normal say?

Show the answerHide the answer

At most 1/9≈11%1/9 \approx 11\%1/9≈11%; a normal gives about 0.27%0.27\%0.27%. Real return distributions typically fall between the two, much nearer the normal in the body and much further from it in the far tail.

Exercise 1.16

More data

How many times as much data would halve the standard error of that Sharpe estimate?

Show the answerHide the answer

Four times as much — sixteen years instead of four — because the error falls with 1/n1/\sqrt n1/n​.

In the interview

When asked "is this result significant?", reach for the standard error before anything else. Naming σ/n\sigma/\sqrt{n}σ/n​ and then asking how independent the observations really are is the answer a researcher gives; quoting a p-value without that step is the answer a student gives.

  • Two Sigma
  • Citadel
  • AQR
  • DE Shaw
Estimation: bias, variance and maximum likelihood →
On this page
  • The law of large numbers
  • The central limit theorem
  • Everything costs four times as much
  • Why the error falls like the square root
  • Worked example
  • When the theorems fail
  • Worked example — how fast the CLT arrives
  • Worked example — the two theorems answer different questions
  • Chebyshev’s inequality
  • Worked example — a guarantee without normality
  • The standard error of a Sharpe ratio
  • Worked example — how precise is four years of Sharpe?
  • When the CLT does not apply
  • Check your understanding
  • Check your understanding — three standard deviations
  • Check your understanding — more data

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.