Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
      • 1Vectors and matrices

        • Rank, trace and determinant: what each one measures
      • 2Linear systems and spaces

        • The four fundamental subspaces, and when Ax = b has a solution
      • 3Orthogonality and projection

        • Orthogonality, QR and OLS as a projection
      • 4Eigenvalues and eigenvectors

        • Eigenvalues, diagonalisation and the spectral theorem
      • 5Definiteness and covariance

        • Definiteness, Cholesky and generating correlated normals
      • 6SVD and dimensionality reduction

        • PCA, covariance matrices and what an eigenvalue is telling you
      • 7Matrix calculus

        • Matrix calculus: the identities behind OLS, ridge and portfolios
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Linear algebra
  4. /Orthogonality and projection

Orthogonality, QR and OLS as a projection

LA · Chapter 3·12 min read·Asked at Two Sigma, Citadel, DE Shaw, Jane Street

Assumes The four fundamental subspaces, and when Ax = b has a solution.

After this lesson you should be able to

  • Build a projection matrix and list its properties.
  • Describe Gram–Schmidt and the QR decomposition it produces.
  • Say why QR is used to solve least squares rather than the normal equations.

Projection is the geometric content of least squares, and orthogonalisation is how it is computed. Understanding both turns a list of regression facts — orthogonal residuals, monotone R2R^2R2, the meaning of "controlling for" — into consequences of one picture.

Equation 3.1

The projection matrix

Projects any vector onto the column space of AAA. Everything about least squares follows from its two defining properties.

P=A(A⊤A)−1A⊤P = A(A^\top A)^{-1}A^\topP=A(A⊤A)−1A⊤
P2=PP^2 = PP2=P
Idempotent — projecting twice is projecting once.
P⊤=PP^\top = PP⊤=P
Symmetric, which is what makes it an *orthogonal* projection.

Proposition 3.2

What the picture gives you

The residual y−Pyy - Pyy−Py is orthogonal to the column space, because the closest point is reached by dropping a perpendicular. That single fact yields: residuals sum to zero when there is an intercept, residuals are uncorrelated with fitted values, and adding a regressor enlarges the space so the residual can only shrink.

Holds when

  • Eigenvalues of a projection are 0 and 1, so tr(P)=rank(P)\mathrm{tr}(P) = \mathrm{rank}(P)tr(P)=rank(P) counts the dimensions projected onto.
  • I−PI - PI−P is also a projection — onto the orthogonal complement — which is where the residuals live.
  • Pythagoras applies: ∥y∥2=∥Py∥2+∥(I−P)y∥2\|y\|^2 = \|Py\|^2 + \|(I-P)y\|^2∥y∥2=∥Py∥2+∥(I−P)y∥2, which *is* the analysis of variance.

Derivation 3.3

Gram–Schmidt and QR

Turn a set of vectors into an orthonormal basis by removing, from each, its component along the ones already processed.

  1. u1=a1,q1=u1/∥u1∥u_1 = a_1, \qquad q_1 = u_1/\|u_1\|u1​=a1​,q1​=u1​/∥u1​∥

    Normalise the first.

  2. uk=ak−∑j<k⟨ak,qj⟩qju_k = a_k - \sum_{j<k}\langle a_k, q_j\rangle q_juk​=ak​−j<k∑​⟨ak​,qj​⟩qj​

    Subtract the part already explained — which is residualising.

  3. A=QR,Q orthonormal, R upper triangularA = QR, \quad Q \text{ orthonormal}, \ R \text{ upper triangular}A=QR,Q orthonormal, R upper triangular
β^=R−1Q⊤y\hat\beta = R^{-1}Q^\top yβ^​=R−1Q⊤y

Why not just invert X⊤XX^\top XX⊤X. Forming X⊤XX^\top XX⊤X squares the condition number, so a design matrix that was merely awkward becomes numerically hopeless: you can lose twice as many digits of precision as the problem itself requires. QR works on XXX directly and never forms the product, so it keeps the original conditioning. On well-conditioned data the two agree; on collinear data, which is the case where you most need the answer, the normal equations can return coefficients that are simply wrong. Every serious least-squares routine uses QR or SVD for this reason.

Example 3.4

Project y=(3,4)y = (3, 4)y=(3,4) onto the line spanned by a=(1,0)a = (1, 0)a=(1,0), and give the residual.

Show the worked solutionHide the worked solution

Worked solution

  1. Formula
    Py=a⊤ya⊤a aPy = \frac{a^\top y}{a^\top a}\,aPy=a⊤aa⊤y​a
  2. Substitute
    a⊤y=3,a⊤a=1a^\top y = 3, \quad a^\top a = 1a⊤y=3,a⊤a=1
  3. Solve
    Py=3(1,0)=(3,0)Py = 3(1, 0) = (3, 0)Py=3(1,0)=(3,0)
  4. y−Py=(0,4)y - Py = (0, 4)y−Py=(0,4)
  5. Answer
    fitted (3,0), residual (0,4)\text{fitted } (3,0), \text{ residual } (0,4)fitted (3,0), residual (0,4)

Sanity check. The residual is orthogonal to aaa, as it must be: (0,4)⋅(1,0)=0(0,4)\cdot(1,0) = 0(0,4)⋅(1,0)=0. And Pythagoras holds — 9+16=25=∥y∥29 + 16 = 25 = \|y\|^29+16=25=∥y∥2 — which is the analysis of variance in two dimensions.

The rest of this lesson is in Premium

You have read the opening. 9 more sections follow, including 3 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← The four fundamental subspaces, and when Ax = b has a solutionEigenvalues, diagonalisation and the spectral theorem →
On this page
  • The projection matrix
  • What the picture gives you
  • Gram–Schmidt and QR
  • Worked example

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.