Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
      • 1Ordinary least squares

        • OLS from three angles
      • 2Assumptions and properties

        • Gauss–Markov: what each assumption buys and what breaks it
      • 3Goodness of fit

        • Goodness of fit: R squared, F tests and information criteria
      • 4Diagnostics

        • Diagnostics: reading residuals, leverage and influence
      • 5Identification

        • Identification: endogeneity, instruments and panel methods
      • 6Regression brainteasers

        • The regression questions firms actually ask
      • 7Regularisation

        • Regularisation: ridge, lasso and choosing lambda
      • 8Generalised models

        • Generalised models: logistic, Poisson and quantile regression
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Regression and econometrics
  4. /Regression brainteasers

The regression questions firms actually ask

REG · Chapter 6·14 min read·Asked at Two Sigma, Citadel, DE Shaw, AQR

Assumes OLS from three angles.

After this lesson you should be able to

  • Explain regression to the mean without invoking a force.
  • Sign an omitted-variable bias from the two correlations involved.
  • Say what adding noise to xxx does, and why it differs from noise in yyy.

There is a short list of regression questions that recur across research interviews, and they test the same thing: whether you reason from the algebra or from a memorised story. Each one below has a one-line answer and a follow-up that catches people who learned only the line.

Definition 6.1

Regression to the mean

Regression to the mean, E[y∣x]−μy=ρ σyσx(x−μx)\mathbb{E}[y \mid x] - \mu_y = \rho\,\frac{\sigma_y}{\sigma_x}(x - \mu_x)E[y∣x]−μy​=ρσx​σy​​(x−μx​) — When xxx and yyy are imperfectly correlated and measured in the same units, an observation far from the mean in xxx is expected to be closer to the mean in yyy — by a factor of exactly ρ\rhoρ. It is not a force pulling things back; it is what imperfect correlation means. The tallest fathers have sons who are tall but less so, and last year’s best fund is expected to be above average this year but not top.

Common trap. Treating regression to the mean as a mechanism — "the market corrects", "talent normalises". If that were a force it would act in one direction in time, and it does not: the sons of tall fathers are shorter, and the fathers of tall sons are also shorter. Instead. Say it is symmetric and therefore not causal. The symmetry is the cleanest way to demonstrate you understand it, and it kills the "so there is a mean-reverting force" follow-up.

-2.502.5-2.502.5Prediction, ρ = 0.6If y simply matched xStandardised xPredicted standardised y
Figure 6.2 · Regression to the mean is not a force. The best prediction of yyy is ρ\rhoρ standard deviations for every one of xxx, so predictions are always flatter than the data. Nothing is pulling anything back to average: it is what "correlation below one" means, and it is why the same effect appears whichever variable you condition on.

Proposition 6.3

Reverse regression

Regress yyy on xxx and you get ρσy/σx\rho\sigma_y/\sigma_xρσy​/σx​; regress xxx on yyy and you get ρσx/σy\rho\sigma_x/\sigma_yρσx​/σy​. The product is ρ2\rho^2ρ2, so the two slopes are reciprocals only when ∣ρ∣=1|\rho| = 1∣ρ∣=1. Plotted on the same axes the two lines both pass through the means and pinch together as the correlation rises.

Holds when

  • If σx=σy\sigma_x = \sigma_yσx​=σy​, both slopes equal ρ\rhoρ, which makes the regression-to-the-mean statement visible directly.
  • Neither line is "the relationship" — each answers a different conditional expectation.

The rest of this lesson is in Premium

You have read the opening. 12 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Identification: endogeneity, instruments and panel methodsRegularisation: ridge, lasso and choosing lambda →
On this page
  • Regression to the mean
  • Regression to the mean is not a force
  • Reverse regression

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.