Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
      • 1Ordinary least squares

        • OLS from three angles
      • 2Assumptions and properties

        • Gauss–Markov: what each assumption buys and what breaks it
      • 3Goodness of fit

        • Goodness of fit: R squared, F tests and information criteria
      • 4Diagnostics

        • Diagnostics: reading residuals, leverage and influence
      • 5Identification

        • Identification: endogeneity, instruments and panel methods
      • 6Regression brainteasers

        • The regression questions firms actually ask
      • 7Regularisation

        • Regularisation: ridge, lasso and choosing lambda
      • 8Generalised models

        • Generalised models: logistic, Poisson and quantile regression
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Regression and econometrics
  4. /Identification

Identification: endogeneity, instruments and panel methods

REG · Chapter 5·13 min read·Asked at Two Sigma, QuantCo, Citadel, AQR

Assumes The regression questions firms actually ask.

After this lesson you should be able to

  • Name the three sources of endogeneity and recognise each.
  • State the two conditions a valid instrument must satisfy.
  • Say what fixed effects and difference-in-differences each remove.

A coefficient is *identified* when the data can distinguish it from the alternatives. Regression alone identifies a causal effect only when the regressor is as good as randomly assigned, and it usually is not — so the field is a catalogue of designs for recovering identification from observational data.

SourceMechanismExample
Omitted variableSomething affects both xxx and yyyAbility affects both schooling and wages
Measurement errorxxx is observed with noiseA proxy for a fundamental, attenuating the coefficient
Simultaneityyyy also causes xxxPrice and quantity determined together
Table 5.1 · Three ways to break exogeneity. All three produce the same symptom — the regressor is correlated with the error — and all three are invisible in the residuals. Which one you have determines the fix.

Definition 5.2

Instrumental variables

A valid instrument, Cov(z,x)≠0andCov(z,ε)=0\mathrm{Cov}(z, x) \ne 0 \quad \text{and} \quad \mathrm{Cov}(z, \varepsilon) = 0Cov(z,x)=0andCov(z,ε)=0 — An instrument zzz moves the regressor without affecting the outcome through any other channel. The first condition — relevance — is testable from the first stage. The second — exclusion — is not testable at all, and must be argued from how the world works. Almost every dispute about an instrumental-variables study is about the second.

Derivation 5.3

Two-stage least squares

Use only the part of xxx that the instrument explains.

  1. x^=fitted values from regressing x on z\hat{x} = \text{fitted values from regressing } x \text{ on } zx^=fitted values from regressing x on z

    The first stage isolates the exogenous variation.

  2. y=βx^+uy = \beta\hat{x} + uy=βx^+u

    The second stage uses only that part, which is uncorrelated with the error by construction.

β^IV=Cov(z,y)Cov(z,x)\hat\beta_{\text{IV}} = \frac{\mathrm{Cov}(z, y)}{\mathrm{Cov}(z, x)}β^​IV​=Cov(z,x)Cov(z,y)​
25010000.40.8IV standard errorFirst-stage F statisticStandard error of the IV estimate
Figure 5.4 · What a weak instrument costs. The curve is steep exactly where instruments usually live. The rule of thumb is F>10F > 10F>10, and the reason is visible: below it the standard error is not merely large, it is rising fast enough that the estimate is dominated by whatever bias the instrument still carries.

Proposition 5.5

Weak instruments are worse than none

When Cov(z,x)\mathrm{Cov}(z,x)Cov(z,x) is small, the instrumental-variables estimator divides by a number near zero: it becomes wildly variable and, worse, biased *toward* the OLS estimate it was supposed to correct. A weak instrument therefore reproduces the bias you were trying to remove while adding enormous variance.

Holds when

  • First-stage F below about 10 is the conventional warning sign.
  • The bias of IV is roughly the OLS bias divided by the first-stage F, so a weak instrument fixes almost nothing.
  • Adding more weak instruments makes it worse, not better.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Diagnostics: reading residuals, leverage and influenceThe regression questions firms actually ask →
On this page
  • Three ways to break exogeneity
  • Instrumental variables
  • Two-stage least squares
  • What a weak instrument costs
  • Weak instruments are worse than none

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.