Skip to content
QuantMax
QuantMax
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • GAMEGames, decision theory and puzzles
    • MMMarket making
    • MKTMarkets and products

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Formula reference

Regression and econometrics

8 lessons · 15 equations. Each lesson below gives its formulas and key rules; open the lesson for the full explanation.

OLS from three angles

The model

y=Xβ+ε,E[ε∣X]=0y = X\beta + \varepsilon, \qquad \mathbb{E}[\varepsilon \mid X] = 0y=Xβ+ε,E[ε∣X]=0

Linear in the parameters — not necessarily in the variables, since x2x^2x2 and log⁡x\log xlogx are perfectly acceptable columns.

The one-regressor case

β^=Cov(x,y)Var(x)=ρ σyσx\hat{\beta} = \frac{\mathrm{Cov}(x, y)}{\mathrm{Var}(x)} = \rho\,\frac{\sigma_y}{\sigma_x}β^​=Var(x)Cov(x,y)​=ρσx​σy​​

Worth memorising in the second form: a beta is a correlation scaled by the ratio of standard deviations.

The variance of the estimator

Var⁡(β^∣X)=σ2(X⊤X)−1\operatorname{Var}(\hat\beta \mid X) = \sigma^2 (X^\top X)^{-1}Var(β^​∣X)=σ2(X⊤X)−1

Under homoskedastic, uncorrelated errors. The diagonal gives each coefficient’s squared standard error; the off-diagonals say how their estimation errors move together. Correlated regressors make X⊤XX^\top XX⊤X nearly singular and blow up the diagonal of its inverse.

Remember

  • β^=(X⊤X)−1X⊤y\hat{\beta} = (X^{\top}X)^{-1}X^{\top}yβ^​=(X⊤X)−1X⊤y, and the residual is orthogonal to every regressor.

Gauss–Markov: what each assumption buys and what breaks it

The sandwich estimator

Var⁡^(β^)=(X⊤X)−1(∑ie^i2 xixi⊤)(X⊤X)−1\widehat{\operatorname{Var}}(\hat\beta) = (X^\top X)^{-1}\Big(\sum_i \hat e_i^2\,x_ix_i^\top\Big)(X^\top X)^{-1}Var(β^​)=(X⊤X)−1(i∑​e^i2​xi​xi⊤​)(X⊤X)−1

White’s heteroskedasticity-robust covariance: the "bread" is the usual (X⊤X)−1(X^\top X)^{-1}(X⊤X)−1 and the "meat" uses each observation’s own squared residual. Newey–West adds weighted cross-products of nearby residuals to handle autocorrelation; clustered versions sum within groups.

Remember

  • Exogeneity and linearity affect the estimate; homoskedasticity and independence affect only inference.

Goodness of fit: R squared, F tests and information criteria

R squared and its adjustment

R2=1−RSSTSS,Rˉ2=1−RSS/(n−k−1)TSS/(n−1)R^2 = 1 - \frac{\mathrm{RSS}}{\mathrm{TSS}}, \qquad \bar{R}^2 = 1 - \frac{\mathrm{RSS}/(n-k-1)}{\mathrm{TSS}/(n-1)}R2=1−TSSRSS​,Rˉ2=1−TSS/(n−1)RSS/(n−k−1)​

The share of variance explained, and the version that charges for parameters.

The F test for nested models

F=(RSSr−RSSu)/qRSSu/(n−k−1)F = \frac{(\mathrm{RSS}_r - \mathrm{RSS}_u)/q}{\mathrm{RSS}_u/(n - k - 1)}F=RSSu​/(n−k−1)(RSSr​−RSSu​)/q​

Tests whether the qqq extra regressors in the unrestricted model add anything jointly.

What one more regressor adds

Rnew2=Rold2+ryxk⋅rest2 (1−Rold2)R^2_{\text{new}} = R^2_{\text{old}} + r^2_{y x_{k} \cdot \text{rest}}\,\big(1 - R^2_{\text{old}}\big)Rnew2​=Rold2​+ryxk​⋅rest2​(1−Rold2​)

A new regressor explains a share of what was left unexplained, equal to its squared partial correlation with yyy given the existing regressors. It can only raise R2R^2R2, and it raises it by little when the model already explains most of the variance.

Remember

  • R2R^2R2 never falls when you add a regressor; adjusted R2R^2R2 can.

Diagnostics: reading residuals, leverage and influence

Cook’s distance

Di=ei2(k+1)s2⋅hii(1−hii)2D_i = \frac{e_i^2}{(k+1)s^2}\cdot\frac{h_{ii}}{(1-h_{ii})^2}Di​=(k+1)s2ei2​​⋅(1−hii​)2hii​​

How much the fitted values move when observation iii is deleted — residual size times leverage.

Studentised residuals

ti=e^iσ^(i)1−hiit_i = \frac{\hat e_i}{\hat\sigma_{(i)}\sqrt{1 - h_{ii}}}ti​=σ^(i)​1−hii​​e^i​​

A residual scaled by its own standard error, estimated without observation iii. High-leverage points have small raw residuals because they pull the line toward themselves; studentising corrects for that, so outliers are compared on a common scale.

Remember

  • Leverage is unusual in XXX; an outlier is unusual in yyy; influence needs both.

Identification: endogeneity, instruments and panel methods

The Wald estimator

β^IV=yˉz=1−yˉz=0xˉz=1−xˉz=0\hat\beta_{IV} = \frac{\bar y_{z=1} - \bar y_{z=0}}{\bar x_{z=1} - \bar x_{z=0}}β^​IV​=xˉz=1​−xˉz=0​yˉ​z=1​−yˉ​z=0​​

With a binary instrument, IV is the effect of the instrument on the outcome divided by its effect on the treatment. It scales up the "intention-to-treat" effect by how much the instrument actually moved the treatment, and it estimates the effect for those whose treatment the instrument changed.

Remember

  • Endogeneity comes from omitted variables, measurement error or simultaneity.

The regression questions firms actually ask

Omitted-variable bias

β^1 → β1+β2 δ,δ=Cov(x1,x2)Var(x1)\hat{\beta}_1 \ \to\ \beta_1 + \beta_2\,\delta, \qquad \delta = \frac{\mathrm{Cov}(x_1, x_2)}{\mathrm{Var}(x_1)}β^​1​ → β1​+β2​δ,δ=Var(x1​)Cov(x1​,x2​)​

Leaving out x2x_2x2​ biases the coefficient on x1x_1x1​ by the true effect of x2x_2x2​ times the regression of x2x_2x2​ on x1x_1x1​.

Remember

  • Regression to the mean is symmetric and therefore not a mechanism.

Regularisation: ridge, lasso and choosing lambda

The two penalties

β^ridge=arg⁡min⁡∥y−Xβ∥2+λ∥β∥22,β^lasso=arg⁡min⁡∥y−Xβ∥2+λ∥β∥1\hat\beta_{\text{ridge}} = \arg\min \|y - X\beta\|^2 + \lambda\|\beta\|_2^2, \qquad \hat\beta_{\text{lasso}} = \arg\min \|y - X\beta\|^2 + \lambda\|\beta\|_1β^​ridge​=argmin∥y−Xβ∥2+λ∥β∥22​,β^​lasso​=argmin∥y−Xβ∥2+λ∥β∥1​

The same least-squares fit plus a penalty on the size of the coefficients — squared for ridge, absolute for lasso.

The elastic net

min⁡β 12∥y−Xβ∥2+λ(α∥β∥1+1−α2∥β∥22)\min_\beta\ \tfrac{1}{2}\lVert y - X\beta\rVert^2 + \lambda\Big(\alpha\lVert\beta\rVert_1 + \tfrac{1 - \alpha}{2}\lVert\beta\rVert_2^2\Big)βmin​ 21​∥y−Xβ∥2+λ(α∥β∥1​+21−α​∥β∥22​)

A blend of the two penalties. The L1L_1L1​ part still sets coefficients to zero; the L2L_2L2​ part makes correlated features share weight rather than having the lasso pick one arbitrarily. α\alphaα between 000 (ridge) and 111 (lasso) is chosen by cross-validation alongside λ\lambdaλ.

Remember

  • Ridge penalises squared coefficients; lasso penalises absolute values.

Generalised models: logistic, Poisson and quantile regression

Logistic regression

ln⁡p1−p=β0+β⊤x⟺p=11+e−(β0+β⊤x)\ln\frac{p}{1-p} = \beta_0 + \beta^\top x \quad \Longleftrightarrow \quad p = \frac{1}{1 + e^{-(\beta_0 + \beta^\top x)}}ln1−pp​=β0​+β⊤x⟺p=1+e−(β0​+β⊤x)1​

Model the log-odds linearly, which maps the whole real line onto (0,1)(0,1)(0,1) so the fitted probability is always valid.

The quantile loss

ρτ(u)=u (τ−1{u<0})\rho_\tau(u) = u\,\big(\tau - \mathbf{1}\{u < 0\}\big)ρτ​(u)=u(τ−1{u<0})

Quantile regression minimises this asymmetric loss. For the median (τ=0.5\tau = 0.5τ=0.5) it is half the absolute error; for a low quantile it charges little for over-predicting and a lot for under-predicting, which pushes the fit down to the chosen quantile.

Remember

  • Logistic models the log-odds; eβe^\betaeβ is the odds multiplier.

Detailed formula cards

  • The regression slope
  • Omitted-variable bias

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.