Estimation: bias, variance and maximum likelihood
STAT · Chapter 212 min readAsked at Two Sigma, Citadel, DE Shaw, QuantCo
Assumes The law of large numbers and the central limit theorem.
After this lesson you should be able to
- Decompose mean squared error into bias and variance.
- Derive a maximum likelihood estimator for the standard cases.
- Say when a biased estimator is the better choice.
An estimator is a recipe, and judging one means asking two separate questions: does it point at the right answer on average, and how much does it move about? Almost every practical choice in estimation is a trade between those two, and in finance the trade usually favours accepting bias.
Equation 2.1
The decomposition
Squared error splits cleanly into a systematic part and a noisy part. Minimising the total is what you actually want; unbiasedness constrains only the first term.
- How far off you are on average.
- How much the estimate moves from sample to sample.
| Property | Statement | Why it matters |
|---|---|---|
| Unbiased | Right on average — says nothing about any one sample | |
| Consistent | as | Enough data eventually gets you there |
| Efficient | Smallest variance in its class | Least wasteful use of the data you have |
| Sufficient | Captures all the information about | You can throw the rest of the data away |
Derivation 2.3
Maximum likelihood
Choose the parameter that makes the data you saw most probable. The standard recipe is three steps.
The likelihood — a function of , with the data fixed.
Take logs, which turns the product into a sum and never moves the maximum.
| Model | MLE | Note |
|---|---|---|
| Bernoulli | The sample proportion | |
| Normal | and | The variance MLE is biased low |
| Exponential | Reciprocal of the sample mean | |
| Poisson | Mean and variance are the same parameter | |
| Uniform | Biased low, and not found by differentiating |
Proposition 2.5
Why the sample variance has
The MLE divides by and is biased low, because it measures deviations from the sample mean rather than the true mean — and the sample mean sits, by construction, in the middle of the data you have. Dividing by corrects exactly for that one degree of freedom spent estimating the mean.
Holds when
- The correction matters at small and is negligible beyond a few hundred observations.
- The unbiased *variance* estimator does not give an unbiased *standard deviation*, because the square root is not linear.
Why bias is often the right trade. Unbiasedness has intuitive appeal and no claim on optimality. If a biased estimator has much smaller variance, its mean squared error is lower and it is simply better. That is the entire case for shrinkage — pulling a noisy estimate toward a structured prior — and in finance the case is overwhelming, because the quantities being estimated have tiny signal-to-noise. A shrunk covariance matrix is biased and enormously more useful than the sample one; the same goes for ridge regression and for shrinking a mean return toward zero.
The rest of this lesson is in Premium
You have read the opening. 11 more sections follow, including 5 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.