The law of large numbers and the central limit theorem
STAT · Chapter 112 min readAsked at Two Sigma, Citadel, DE Shaw, Jane Street
After this lesson you should be able to
- State what each theorem says, and what it does not.
- Use the rate to size a sample or a backtest.
- Name the conditions under which both fail on financial data.
Two theorems carry most of applied statistics. The law of large numbers says an average converges to the truth; the central limit theorem says how fast and in what shape. Almost every question about sample size, error bars or the significance of a backtest is one of these in disguise.
Definition 1.1
The law of large numbers
LLN, — The sample mean of independent draws with a finite mean converges to that mean. The weak form says it converges in probability, the strong form almost surely; for interview purposes the distinction rarely matters, but knowing there is one does.
Equation 1.2
The central limit theorem
The standardised sample mean converges in distribution to a standard normal, whatever the shape of the underlying distribution — provided the variance is finite.
- The standard error: the standard deviation of the sample mean.
- Convergence in distribution — the shape converges, not the individual values.
Proposition 1.3
Everything costs four times as much
Error shrinks as , so halving your error bar takes four times the data. This single fact sets the economics of research: an effect you can measure to one decimal place in a year of data needs a century for two. It is also why Monte Carlo is slow and why short backtests prove very little.
Holds when
- To cut a confidence interval in half, quadruple the sample.
- Monte Carlo error is regardless of dimension, which is why it beats quadrature in high dimensions and loses in low ones.
Example 1.5
A strategy has a true Sharpe ratio of . Roughly how many years of daily returns do you need before the estimated Sharpe is significant at two standard errors?
Show the worked solutionHide the worked solution
Worked solution
- Formula
- Substitute
- Solve
- Answer
Sanity check. A Sharpe of 0.5 would need sixteen years, which is the honest reason weak signals are so hard to validate.
| Violation | What happens | Where it shows up |
|---|---|---|
| Infinite variance | CLT does not apply; the average does not concentrate | Cauchy-like tails, ratios of returns |
| Infinite mean | Even the LLN fails | The St. Petersburg payoff; pure Pareto with |
| Dependence | Effective sample size is far below | Autocorrelated returns, overlapping windows |
| Non-identical distributions | Convergence may still hold, but slowly | Regime change, structural breaks |
Why normality appears at all. A sum forgets the details of its parts. Adding many independent contributions, each small relative to the total, washes out everything except the mean and the variance — the higher moments of the individual terms scale away faster than the sum grows. That is why the normal turns up everywhere, and equally why it stops turning up when a single term can dominate the sum, which is exactly what a fat tail is.
Common trap. Claiming the CLT says the data become normal. It says nothing about the data; it is a statement about the sampling distribution of an average. Daily equity returns are not normal and never become so no matter how many you collect. Instead. Say what converges: the standardised *mean*, not the observations. Averages of returns are much closer to normal than the returns themselves.
Example 1.7
How fast the CLT arrives
Daily returns are heavily skewed. You average of them and then . How much does the skew of the mean fall?
Show the worked solutionHide the worked solution
Worked solution
- Formula
- Substitute
- Solve
- Answer
Sanity check. Skew dies like and kurtosis like , which is why monthly return series look normal and daily ones never do — and why a test calibrated on monthly data misfires on daily.
Example 1.8
The two theorems answer different questions
You toss a fair coin a million times. How close is the proportion to a half, and how many heads should you expect the count to be away from ?
Show the worked solutionHide the worked solution
Worked solution
- Formula
- Substitute
- Solve
- Answer
Sanity check. The fraction gets closer to a half and the absolute number of surplus heads gets further from zero. Both are true and candidates who have only heard of the law of large numbers get the second one wrong.
Equation 1.9
Chebyshev’s inequality
Holds for any distribution with a finite variance — no normality required. It is loose, but it is the guarantee you fall back on when you cannot trust the tails, and it is what proves the weak law of large numbers in one line.
Example 1.10
A guarantee without normality
A sample mean has standard error . Without assuming normality, what is the most you can say about the chance it lands two standard errors or more from the truth?
Show the worked solutionHide the worked solution
Worked solution
- Formula
- Substitute
- Solve
- Answer
Sanity check. Under normality the answer is . The gap between the two is the price of making no assumption about the tails — and heavy-tailed returns sit somewhere between.
Proposition 1.13
When the CLT does not apply
The central limit theorem needs a finite variance. For a Student- with three degrees of freedom the variance is finite and averages do become normal, but slowly; with two or fewer the variance is infinite, and suitably scaled sums converge to a stable, non-normal law instead. Returns with tail exponents near three are why "approximately normal averages" can take far more data than the textbook suggests.
Holds when
- Finite variance: CLT, at rate , slower with heavy tails and skew.
- Infinite variance: no CLT; sums scale like with .
How fast normality arrives. The Berry–Esseen theorem bounds the gap between the distribution of a standardised mean and the normal by a constant times the skewness over . Symmetric data become normal quickly; skewed data — option payoffs, default losses — can need thousands of observations before tail probabilities from a normal approximation can be trusted.
Common trap — averaging ratios. Averaging per-period Sharpe ratios, or averaging returns on different capital bases, and calling it the ratio for the whole. The average of ratios is not the ratio of averages. Instead. Compute the statistic on the pooled data — total mean over total standard deviation — or weight the pieces properly. For returns, average log returns or compound the simple ones.
What you need to know
- LLN: the sample mean converges to the true mean, given a finite mean.
- CLT: the standardised sample mean converges to a standard normal, given a finite variance.
- The standard error is — quartering the error takes sixteen times the data.
- The CLT is about the distribution of an average, not about the data.
- Dependence is the practical killer: it shrinks the effective sample size without shrinking .
Exercise 1.14
You compute a monthly Sharpe ratio from overlapping twelve-month returns sampled every month over ten years. Why is the resulting error bar too small?
Show the answerHide the answer
Consecutive observations share eleven of their twelve months, so they are heavily correlated. You have 120 numbers but roughly ten independent ones, and dividing by rather than understates the standard error by a factor of about three and a half.
Exercise 1.15
Three standard deviations
What does Chebyshev guarantee about a move of at least three standard deviations, and what would a normal say?
Show the answerHide the answer
At most ; a normal gives about . Real return distributions typically fall between the two, much nearer the normal in the body and much further from it in the far tail.
Exercise 1.16
More data
How many times as much data would halve the standard error of that Sharpe estimate?
Show the answerHide the answer
Four times as much — sixteen years instead of four — because the error falls with .
In the interview
When asked "is this result significant?", reach for the standard error before anything else. Naming and then asking how independent the observations really are is the answer a researcher gives; quoting a p-value without that step is the answer a student gives.
- Two Sigma
- Citadel
- AQR
- DE Shaw