Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
      • 1Limit theorems

        • The law of large numbers and the central limit theorem
      • 2Estimation

        • Estimation: bias, variance and maximum likelihood
      • 3Confidence intervals

        • Confidence intervals, and the interval you would trade
      • 4Hypothesis testing

        • Hypothesis testing: errors, power and which test to use
      • 5p-values and multiple testing

        • p-values, p-hacking and the multiple-testing problem
      • 6Resampling

        • Resampling: the bootstrap, permutation tests and where they break
      • 7Bayesian statistics

        • Bayesian statistics: conjugacy, shrinkage and credible intervals
      • 8Experiment design

        • Experiment design: randomisation, peeking and minimum detectable effect
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Statistics and inference
  4. /Estimation

Estimation: bias, variance and maximum likelihood

STAT · Chapter 2·12 min read·Asked at Two Sigma, Citadel, DE Shaw, QuantCo

Assumes The law of large numbers and the central limit theorem.

After this lesson you should be able to

  • Decompose mean squared error into bias and variance.
  • Derive a maximum likelihood estimator for the standard cases.
  • Say when a biased estimator is the better choice.

An estimator is a recipe, and judging one means asking two separate questions: does it point at the right answer on average, and how much does it move about? Almost every practical choice in estimation is a trade between those two, and in finance the trade usually favours accepting bias.

Equation 2.1

The decomposition

Squared error splits cleanly into a systematic part and a noisy part. Minimising the total is what you actually want; unbiasedness constrains only the first term.

MSE(θ^)=E[(θ^−θ)2]=Bias(θ^)2+Var(θ^)\mathrm{MSE}(\hat\theta) = \mathbb{E}\big[(\hat\theta - \theta)^2\big] = \mathrm{Bias}(\hat\theta)^2 + \mathrm{Var}(\hat\theta)MSE(θ^)=E[(θ^−θ)2]=Bias(θ^)2+Var(θ^)
Bias=E[θ^]−θ\mathrm{Bias} = \mathbb{E}[\hat\theta] - \thetaBias=E[θ^]−θ
How far off you are on average.
Var\mathrm{Var}Var
How much the estimate moves from sample to sample.
PropertyStatementWhy it matters
UnbiasedE[θ^]=θ\mathbb{E}[\hat\theta] = \thetaE[θ^]=θRight on average — says nothing about any one sample
Consistentθ^→θ\hat\theta \to \thetaθ^→θ as n→∞n \to \inftyn→∞Enough data eventually gets you there
EfficientSmallest variance in its classLeast wasteful use of the data you have
SufficientCaptures all the information about θ\thetaθYou can throw the rest of the data away
Table 2.2 · What the words mean. Unbiased and consistent are independent properties: an estimator can be unbiased and never converge, or biased at every nnn and still consistent.

Derivation 2.3

Maximum likelihood

Choose the parameter that makes the data you saw most probable. The standard recipe is three steps.

  1. L(θ)=∏if(xi;θ)L(\theta) = \prod_i f(x_i; \theta)L(θ)=i∏​f(xi​;θ)

    The likelihood — a function of θ\thetaθ, with the data fixed.

  2. ℓ(θ)=∑iln⁡f(xi;θ)\ell(\theta) = \sum_i \ln f(x_i; \theta)ℓ(θ)=i∑​lnf(xi​;θ)

    Take logs, which turns the product into a sum and never moves the maximum.

  3. ∂ℓ∂θ=0\frac{\partial \ell}{\partial \theta} = 0∂θ∂ℓ​=0
θ^MLE\hat\theta_{\text{MLE}}θ^MLE​
ModelMLENote
Bernoulli(p)(p)(p)xˉ\bar{x}xˉThe sample proportion
Normal(μ,σ2)(\mu, \sigma^2)(μ,σ2)xˉ\bar{x}xˉ and 1n∑(xi−xˉ)2\frac{1}{n}\sum(x_i - \bar{x})^2n1​∑(xi​−xˉ)2The variance MLE is biased low
Exponential(λ)(\lambda)(λ)1/xˉ1/\bar{x}1/xˉReciprocal of the sample mean
Poisson(λ)(\lambda)(λ)xˉ\bar{x}xˉMean and variance are the same parameter
Uniform(0,θ)(0, \theta)(0,θ)max⁡xi\max x_imaxxi​Biased low, and not found by differentiating
Table 2.4 · The MLEs worth knowing. The last row is the one interviewers use: the likelihood is maximised at a boundary, so calculus finds nothing and you have to reason about the support instead.

Proposition 2.5

Why the sample variance has n−1n-1n−1

The MLE divides by nnn and is biased low, because it measures deviations from the sample mean rather than the true mean — and the sample mean sits, by construction, in the middle of the data you have. Dividing by n−1n-1n−1 corrects exactly for that one degree of freedom spent estimating the mean.

Holds when

  • The correction matters at small nnn and is negligible beyond a few hundred observations.
  • The unbiased *variance* estimator does not give an unbiased *standard deviation*, because the square root is not linear.

Why bias is often the right trade. Unbiasedness has intuitive appeal and no claim on optimality. If a biased estimator has much smaller variance, its mean squared error is lower and it is simply better. That is the entire case for shrinkage — pulling a noisy estimate toward a structured prior — and in finance the case is overwhelming, because the quantities being estimated have tiny signal-to-noise. A shrunk covariance matrix is biased and enormously more useful than the sample one; the same goes for ridge regression and for shrinking a mean return toward zero.

The rest of this lesson is in Premium

You have read the opening. 11 more sections follow, including 5 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← The law of large numbers and the central limit theoremConfidence intervals, and the interval you would trade →
On this page
  • The decomposition
  • What the words mean
  • Maximum likelihood
  • The MLEs worth knowing
  • Why the sample variance has n−1n-1n−1

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.