Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
      • 1Ordinary least squares

        • OLS from three angles
      • 2Assumptions and properties

        • Gauss–Markov: what each assumption buys and what breaks it
      • 3Goodness of fit

        • Goodness of fit: R squared, F tests and information criteria
      • 4Diagnostics

        • Diagnostics: reading residuals, leverage and influence
      • 5Identification

        • Identification: endogeneity, instruments and panel methods
      • 6Regression brainteasers

        • The regression questions firms actually ask
      • 7Regularisation

        • Regularisation: ridge, lasso and choosing lambda
      • 8Generalised models

        • Generalised models: logistic, Poisson and quantile regression
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Regression and econometrics
  4. /Regularisation

Regularisation: ridge, lasso and choosing lambda

REG · Chapter 7·12 min read·Asked at Two Sigma, Citadel, QuantCo, AQR

Assumes Goodness of fit: R squared, F tests and information criteria.

After this lesson you should be able to

  • Write down the ridge and lasso objectives and say what each penalty does.
  • Explain why lasso produces exactly zero coefficients and ridge does not.
  • Choose lambda honestly on a time series.

Regularisation deliberately biases the coefficients toward zero in exchange for a large reduction in variance. On financial data, where the signal is faint and the regressors are correlated, that trade is almost always worth making — which is why a penalised regression routinely beats OLS out of sample.

Equation 7.1

The two penalties

The same least-squares fit plus a penalty on the size of the coefficients — squared for ridge, absolute for lasso.

β^ridge=arg⁡min⁡∥y−Xβ∥2+λ∥β∥22,β^lasso=arg⁡min⁡∥y−Xβ∥2+λ∥β∥1\hat\beta_{\text{ridge}} = \arg\min \|y - X\beta\|^2 + \lambda\|\beta\|_2^2, \qquad \hat\beta_{\text{lasso}} = \arg\min \|y - X\beta\|^2 + \lambda\|\beta\|_1β^​ridge​=argmin∥y−Xβ∥2+λ∥β∥22​,β^​lasso​=argmin∥y−Xβ∥2+λ∥β∥1​
λ\lambdaλ
The strength of the penalty. At zero this is OLS; as it grows everything shrinks toward zero.
∥β∥1\|\beta\|_1∥β∥1​
Sum of absolute values, which is what produces exact zeros.

Proposition 7.2

Ridge has a closed form and always works

The solution is β^=(X⊤X+λI)−1X⊤y\hat\beta = (X^\top X + \lambda I)^{-1}X^\top yβ^​=(X⊤X+λI)−1X⊤y. Adding λ\lambdaλ to the diagonal makes the matrix invertible even when X⊤XX^\top XX⊤X is singular, so ridge works when there are more regressors than observations and OLS does not exist at all.

Holds when

  • Standardise the regressors first, or the penalty charges more for variables measured in small units.
  • Do not penalise the intercept.
  • Ridge has a Bayesian reading: it is the posterior mode under a normal prior on β\betaβ centred at zero.

Why lasso selects and ridge does not. Picture the constraint region each penalty defines. Ridge gives a sphere, lasso a diamond with corners on the axes. The solution is where the least-squares contours first touch that region, and a smooth sphere is almost never touched exactly at an axis — so ridge coefficients get small but stay non-zero. A diamond has corners *on* the axes, and contours meet corners readily, so lasso sets coefficients exactly to zero. Selection is a consequence of the geometry, not an extra feature.

AspectRidgeLassoElastic net
Penalty∥β∥22\|\beta\|_2^2∥β∥22​∥β∥1\|\beta\|_1∥β∥1​A mix of both
Exact zerosNoYesYes
Correlated regressorsSplits weight between themPicks one arbitrarilyKeeps the group together
Closed formYesNoNo
Use whenAll regressors plausibly matterYou want a sparse modelCorrelated groups of features
Table 7.3 · Ridge, lasso, elastic net. The correlated-regressors row decides it in practice. On a factor library where features come in families, lasso picks one member of each at random and elastic net keeps the family, which is far more stable across refits.
012300.51Ridge: β/(1+λ)Lasso: max(β−λ, 0)Penalty λFitted coefficient
Figure 7.4 · Ridge bends, lasso cuts. A coefficient that starts at 111. Ridge shrinks it towards zero and never reaches it, so every regressor stays in the model; lasso subtracts a constant and hits zero at λ=1\lambda = 1λ=1, which is the whole of why one selects variables and the other does not.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← The regression questions firms actually askGeneralised models: logistic, Poisson and quantile regression →
On this page
  • The two penalties
  • Ridge has a closed form and always works
  • Ridge, lasso, elastic net
  • Ridge bends, lasso cuts

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.