Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
      • 1Ordinary least squares

        • OLS from three angles
      • 2Assumptions and properties

        • Gauss–Markov: what each assumption buys and what breaks it
      • 3Goodness of fit

        • Goodness of fit: R squared, F tests and information criteria
      • 4Diagnostics

        • Diagnostics: reading residuals, leverage and influence
      • 5Identification

        • Identification: endogeneity, instruments and panel methods
      • 6Regression brainteasers

        • The regression questions firms actually ask
      • 7Regularisation

        • Regularisation: ridge, lasso and choosing lambda
      • 8Generalised models

        • Generalised models: logistic, Poisson and quantile regression
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Regression and econometrics
  4. /Generalised models

Generalised models: logistic, Poisson and quantile regression

REG · Chapter 8·12 min read·Asked at Two Sigma, QuantCo, Citadel, Point72

Assumes Regularisation: ridge, lasso and choosing lambda.

After this lesson you should be able to

  • Interpret a logistic coefficient as a log-odds and convert it.
  • Say when a Poisson model is the right one and how to check it.
  • Explain what quantile regression estimates that OLS does not.

Linear regression models a conditional mean on an unbounded scale. When the outcome is a probability, a count or a tail, that is the wrong target — and each generalised model is the same linear predictor pushed through a function that makes the range come out right.

Equation 8.1

Logistic regression

Model the log-odds linearly, which maps the whole real line onto (0,1)(0,1)(0,1) so the fitted probability is always valid.

ln⁡p1−p=β0+β⊤x⟺p=11+e−(β0+β⊤x)\ln\frac{p}{1-p} = \beta_0 + \beta^\top x \quad \Longleftrightarrow \quad p = \frac{1}{1 + e^{-(\beta_0 + \beta^\top x)}}ln1−pp​=β0​+β⊤x⟺p=1+e−(β0​+β⊤x)1​
p/(1−p)p/(1-p)p/(1−p)
The odds. A coefficient of β\betaβ multiplies the odds by eβe^\betaeβ per unit.
β0\beta_0β0​
Log-odds at x=0x = 0x=0, which is rarely a meaningful point unless you centre.
-60600.511/(1+e⁻ˣᵝ)Linear predictor xβProbability
Figure 8.2 · The link, and where the slope lives. The coefficient is constant in log-odds and anything but constant in probability: the same one-unit move is worth 252525 percentage points at the middle and almost nothing in the tails. Quoting a logistic coefficient as "an effect on probability" without saying where is the standard error.

Proposition 8.3

Reading a logistic coefficient

A coefficient of 0.70.70.7 means the odds multiply by e0.7≈2e^{0.7} \approx 2e0.7≈2 per unit increase. Near a probability of 0.50.50.5 the effect on the *probability* is about β/4\beta/4β/4 per unit — the derivative of the logistic at its midpoint is one quarter — which is a fast way to translate a coefficient into something interpretable.

Holds when

  • The divide-by-four rule is an upper bound: the effect on probability is smaller away from 0.50.50.5.
  • Odds ratios are not probability ratios, and conflating them overstates effects on common outcomes.
  • No R2R^2R2 exists; use log-likelihood, AUC or a calibration plot instead.

Example 8.4

A logistic model of whether a trade is profitable gives a coefficient of 0.40.40.4 on a signal. The base rate is 50%50\%50%. What does a one-unit increase do?

Show the worked solutionHide the worked solution

Worked solution

  1. Formula
    odds×eβ,Δp≈β/4 near p=0.5\text{odds} \times e^{\beta}, \qquad \Delta p \approx \beta/4 \text{ near } p = 0.5odds×eβ,Δp≈β/4 near p=0.5
  2. Substitute
    e0.4=1.49e^{0.4} = 1.49e0.4=1.49
  3. Solve
    odds: 1.0→1.49⇒p=1.492.49=0.598\text{odds: } 1.0 \to 1.49 \Rightarrow p = \frac{1.49}{2.49} = 0.598odds: 1.0→1.49⇒p=2.491.49​=0.598
  4. rule of thumb: 0.5+0.4/4=0.60\text{rule of thumb: } 0.5 + 0.4/4 = 0.60rule of thumb: 0.5+0.4/4=0.60
  5. Answer
    about 60%, up from 50%\text{about } 60\%, \text{ up from } 50\%about 60%, up from 50%

Sanity check. The divide-by-four approximation gives 0.600.600.60 against an exact 0.5980.5980.598 — close enough to quote, and it fails gracefully by overstating as you move away from the midpoint.

The rest of this lesson is in Premium

You have read the opening. 12 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Regularisation: ridge, lasso and choosing lambdaBack to Regression and econometrics →
On this page
  • Logistic regression
  • The link, and where the slope lives
  • Reading a logistic coefficient
  • Worked example

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.