Regularisation: ridge, lasso and choosing lambda
REG · Chapter 712 min readAsked at Two Sigma, Citadel, QuantCo, AQR
Assumes Goodness of fit: R squared, F tests and information criteria.
After this lesson you should be able to
- Write down the ridge and lasso objectives and say what each penalty does.
- Explain why lasso produces exactly zero coefficients and ridge does not.
- Choose lambda honestly on a time series.
Regularisation deliberately biases the coefficients toward zero in exchange for a large reduction in variance. On financial data, where the signal is faint and the regressors are correlated, that trade is almost always worth making — which is why a penalised regression routinely beats OLS out of sample.
Equation 7.1
The two penalties
The same least-squares fit plus a penalty on the size of the coefficients — squared for ridge, absolute for lasso.
- The strength of the penalty. At zero this is OLS; as it grows everything shrinks toward zero.
- Sum of absolute values, which is what produces exact zeros.
Proposition 7.2
Ridge has a closed form and always works
The solution is . Adding to the diagonal makes the matrix invertible even when is singular, so ridge works when there are more regressors than observations and OLS does not exist at all.
Holds when
- Standardise the regressors first, or the penalty charges more for variables measured in small units.
- Do not penalise the intercept.
- Ridge has a Bayesian reading: it is the posterior mode under a normal prior on centred at zero.
Why lasso selects and ridge does not. Picture the constraint region each penalty defines. Ridge gives a sphere, lasso a diamond with corners on the axes. The solution is where the least-squares contours first touch that region, and a smooth sphere is almost never touched exactly at an axis — so ridge coefficients get small but stay non-zero. A diamond has corners *on* the axes, and contours meet corners readily, so lasso sets coefficients exactly to zero. Selection is a consequence of the geometry, not an extra feature.
| Aspect | Ridge | Lasso | Elastic net |
|---|---|---|---|
| Penalty | A mix of both | ||
| Exact zeros | No | Yes | Yes |
| Correlated regressors | Splits weight between them | Picks one arbitrarily | Keeps the group together |
| Closed form | Yes | No | No |
| Use when | All regressors plausibly matter | You want a sparse model | Correlated groups of features |
The rest of this lesson is in Premium
You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.