Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
      • 1Ordinary least squares

        • OLS from three angles
      • 2Assumptions and properties

        • Gauss–Markov: what each assumption buys and what breaks it
      • 3Goodness of fit

        • Goodness of fit: R squared, F tests and information criteria
      • 4Diagnostics

        • Diagnostics: reading residuals, leverage and influence
      • 5Identification

        • Identification: endogeneity, instruments and panel methods
      • 6Regression brainteasers

        • The regression questions firms actually ask
      • 7Regularisation

        • Regularisation: ridge, lasso and choosing lambda
      • 8Generalised models

        • Generalised models: logistic, Poisson and quantile regression
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Regression and econometrics
  4. /Ordinary least squares

OLS from three angles

REG · Chapter 1·13 min read·Asked at Two Sigma, Citadel, DE Shaw, AQR

After this lesson you should be able to

  • Derive the OLS estimator by calculus and by projection.
  • Interpret a coefficient in a multiple regression precisely.
  • Say what each Gauss–Markov assumption buys you.

Least squares is the workhorse, and interviewers probe whether you understand it or merely call it. Three views of the same object — minimise a sum of squares, project onto a column space, take a ratio of covariance to variance — and each answers a different follow-up.

Equation 1.1

The model

Linear in the parameters — not necessarily in the variables, since x2x^2x2 and log⁡x\log xlogx are perfectly acceptable columns.

y=Xβ+ε,E[ε∣X]=0y = X\beta + \varepsilon, \qquad \mathbb{E}[\varepsilon \mid X] = 0y=Xβ+ε,E[ε∣X]=0
XXX
The n×kn \times kn×k design matrix, including a column of ones for the intercept.
ε\varepsilonε
The error, assumed mean-zero conditional on the regressors.

Derivation 1.2

The normal equations

Minimise the sum of squared residuals with respect to β\betaβ.

  1. S(β)=(y−Xβ)⊤(y−Xβ)S(\beta) = (y - X\beta)^{\top}(y - X\beta)S(β)=(y−Xβ)⊤(y−Xβ)
  2. ∂S∂β=−2X⊤(y−Xβ)=0\frac{\partial S}{\partial \beta} = -2X^{\top}(y - X\beta) = 0∂β∂S​=−2X⊤(y−Xβ)=0

    Setting the gradient to zero gives the normal equations.

  3. X⊤Xβ^=X⊤yX^{\top}X\hat{\beta} = X^{\top}yX⊤Xβ^​=X⊤y
β^=(X⊤X)−1X⊤y\hat{\beta} = (X^{\top}X)^{-1}X^{\top}yβ^​=(X⊤X)−1X⊤y

The projection view. The normal equations say X⊤(y−Xβ^)=0X^{\top}(y - X\hat{\beta}) = 0X⊤(y−Xβ^​)=0: the residual is orthogonal to every column of XXX. So y^\hat{y}y^​ is the point in the column space of XXX closest to yyy, and H=X(X⊤X)−1X⊤H = X(X^{\top}X)^{-1}X^{\top}H=X(X⊤X)−1X⊤ is the projection matrix onto that space. Everything mechanical about OLS follows from this picture — residuals sum to zero when there is an intercept, adding a regressor can never increase the residual sum of squares, and fitted values are uncorrelated with residuals.

Equation 1.3

The one-regressor case

Worth memorising in the second form: a beta is a correlation scaled by the ratio of standard deviations.

β^=Cov(x,y)Var(x)=ρ σyσx\hat{\beta} = \frac{\mathrm{Cov}(x, y)}{\mathrm{Var}(x)} = \rho\,\frac{\sigma_y}{\sigma_x}β^​=Var(x)Cov(x,y)​=ρσx​σy​​
ρ\rhoρ
The correlation between xxx and yyy.
01242444RSS(b)Candidate slope bResidual sum of squares
Figure 1.4 · Least squares is one parabola. With one regressor the objective is a parabola in the slope, so it has exactly one minimum and a closed form. Everything OLS is — the normal equations, the projection, the uniqueness — is this shape in as many dimensions as you have regressors.

Proposition 1.5

What a coefficient means

In a multiple regression, βj\beta_jβj​ is the expected change in yyy for a one-unit change in xjx_jxj​ *holding the other regressors fixed*. That clause is the whole content of the estimate, and by the Frisch–Waugh–Lovell theorem it is literally true: βj\beta_jβj​ equals the simple regression coefficient of the residualised yyy on the residualised xjx_jxj​.

Holds when

  • Holding fixed is a statistical operation, not a causal one — it is not the same as intervening.
  • If xjx_jxj​ is nearly collinear with the others there is little residual variation left, which is why the standard error explodes.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

Gauss–Markov: what each assumption buys and what breaks it →
On this page
  • The model
  • The normal equations
  • The one-regressor case
  • Least squares is one parabola
  • What a coefficient means

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.