Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
      • 1Ordinary least squares

        • OLS from three angles
      • 2Assumptions and properties

        • Gauss–Markov: what each assumption buys and what breaks it
      • 3Goodness of fit

        • Goodness of fit: R squared, F tests and information criteria
      • 4Diagnostics

        • Diagnostics: reading residuals, leverage and influence
      • 5Identification

        • Identification: endogeneity, instruments and panel methods
      • 6Regression brainteasers

        • The regression questions firms actually ask
      • 7Regularisation

        • Regularisation: ridge, lasso and choosing lambda
      • 8Generalised models

        • Generalised models: logistic, Poisson and quantile regression
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Regression and econometrics
  4. /Diagnostics

Diagnostics: reading residuals, leverage and influence

REG · Chapter 4·12 min read·Asked at Two Sigma, Citadel, QuantCo, AQR

Assumes Gauss–Markov: what each assumption buys and what breaks it.

After this lesson you should be able to

  • Diagnose a model from its residual plots.
  • Distinguish leverage, outlier and influence.
  • Say what each diagnostic can and cannot detect.

A regression output is four numbers and a great deal of hidden behaviour. The diagnostics exist to surface that behaviour, and each one detects a specific failure — which means knowing what a plot *cannot* reveal is as useful as knowing what it can.

PlotLooks forA problem looks like
Residuals against fittedNon-linearity, heteroskedasticityCurvature, or a widening fan
Residuals against each regressorA missing transformationStructure in one variable only
Q-Q plot of residualsNon-normalityHeavy tails curving away at the ends
Residuals against time or orderAutocorrelation, regime changeRuns of the same sign
Scale–locationHeteroskedasticity specificallyA trend in the spread
Residuals against leverageInfluential pointsA point outside the Cook’s distance contour
Table 4.1 · What each plot shows. None of these can detect endogeneity. Residuals are orthogonal to the regressors by construction, so an omitted variable correlated with a regressor leaves no trace in any plot.

Definition 4.2

Leverage, outlier, influence

Three different things, hii=[X(X⊤X)−1X⊤]iih_{ii} = \big[X(X^\top X)^{-1}X^\top\big]_{ii}hii​=[X(X⊤X)−1X⊤]ii​ — A high-*leverage* point is unusual in XXX — far from the centre of the regressors. An *outlier* has a large residual, unusual in yyy given xxx. A point is *influential* only when it is both: unusual in XXX and not on the line the rest of the data describe. The diagonal of the hat matrix measures leverage, and the rule of thumb is that hiih_{ii}hii​ above 2(k+1)/n2(k+1)/n2(k+1)/n is worth a look.

Why influence needs both. A point in the middle of the data with a large residual barely moves the fit — there is plenty of other data at that xxx to hold the line in place. A point far out in XXX but sitting exactly on the trend also changes nothing; it simply confirms it, with a lot of weight. It is the combination that does damage: a far-out point off the line pivots the whole regression toward itself, and can single-handedly create or destroy a coefficient. Cook’s distance is exactly the product of the two ingredients, which is why it is the statistic that matters.

Equation 4.3

Cook’s distance

How much the fitted values move when observation iii is deleted — residual size times leverage.

Di=ei2(k+1)s2⋅hii(1−hii)2D_i = \frac{e_i^2}{(k+1)s^2}\cdot\frac{h_{ii}}{(1-h_{ii})^2}Di​=(k+1)s2ei2​​⋅(1−hii​)2hii​​
eie_iei​
The residual: how much of an outlier the point is in yyy.
hiih_{ii}hii​
The leverage: how unusual it is in XXX.
02468100246810Without pointWith pointObservation 44Factor exposureResponse
Figure 4.4 · One observation pivots the fitted line. The highlighted observation sits far from the others in both factor exposure and response. Including it pulls the fitted slope from about 0.500.500.50 to 0.750.750.75. Cook’s distance detects this combination of leverage and residual size; the honest report shows the fit with and without the point.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Goodness of fit: R squared, F tests and information criteriaIdentification: endogeneity, instruments and panel methods →
On this page
  • What each plot shows
  • Leverage, outlier, influence
  • Cook’s distance
  • One observation pivots the fitted line

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.