Gauss–Markov: what each assumption buys and what breaks it
REG · Chapter 212 min readAsked at Two Sigma, Citadel, DE Shaw, AQR
Assumes OLS from three angles.
After this lesson you should be able to
- State the Gauss–Markov assumptions and what each one delivers.
- Separate assumptions that affect the estimate from those that affect only inference.
- Say what to do when each one fails.
The Gauss–Markov theorem says OLS is the best linear unbiased estimator under a specific list of conditions. The list is worth knowing not as a recitation but as a diagnostic: when a regression misbehaves, one of these has failed, and which one tells you whether the coefficient is wrong or only the standard error.
| Assumption | Failure affects | Fix |
|---|---|---|
| Linear in parameters | The estimate — you are fitting the wrong thing | Transform variables, or use a different model |
| The estimate — bias | Instruments, controls, panel methods | |
| No perfect collinearity | Identification — no unique solution | Drop a regressor, or regularise |
| Homoskedasticity | Inference only | Robust (White) standard errors |
| No autocorrelation | Inference only | Newey–West / HAC standard errors |
| Normal errors | Small-sample inference only | Nothing — the CLT covers large samples |
Definition 2.2
What BLUE actually claims
Best linear unbiased estimator — Smallest variance among linear, unbiased estimators. Both qualifiers matter: a non-linear estimator may do better, and a biased one frequently does — which is exactly the opening for ridge regression and every other shrinkage method. BLUE is a strong-sounding result about a deliberately narrow class.
Exogeneity is the one that matters. Of the six, only can make your coefficient point at the wrong number, and it is also the only one you cannot test — the residuals are orthogonal to the regressors by construction, so a diagnostic plot can never reveal the failure. It has to be argued from what you know about how the data were generated. Everything else on the list is checkable from the output; this one is checkable only from the world.
Proposition 2.3
Robust standard errors
Heteroskedasticity and autocorrelation leave OLS unbiased and make the usual standard errors wrong — usually too small, so everything looks more significant than it is. White standard errors fix heteroskedasticity; Newey–West additionally handles autocorrelation up to a chosen lag. Neither changes the coefficient.
Holds when
- On panel data, cluster by the unit that shares the shock — usually the firm or the date.
- Clustering by date is the standard correction for cross-sectional return regressions, where every stock shares the market move.
- Robust errors are nearly free, so the sensible default is to use them.
The rest of this lesson is in Premium
You have read the opening. 11 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.