Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies
      • 1A method for open cases

        • A method for open-ended research cases
      • 2Prediction cases

        • Prediction cases: demand, pricing and the traps in each
      • 3Causal and experiment cases

        • Causal cases: treatment effects, confounding and surprising results
      • 4Market-data cases

        • Market-data cases: short-horizon prediction from the order book
      • 5Data quality

        • Data quality: missingness, outliers, regimes and leakage
      • 6Communicating

        • Communicating: structuring an answer and defending an assumption

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Research case studies
  4. /Causal and experiment cases

Causal cases: treatment effects, confounding and surprising results

CASE · Chapter 3·12 min read·Asked at QuantCo, Two Sigma, Citadel, Point72

Assumes Prediction cases: demand, pricing and the traps in each.

After this lesson you should be able to

  • State what a treatment effect is and why it needs a counterfactual.
  • Recognise selection bias, Simpson’s paradox and survivorship in a case.
  • Say what you would do when randomisation is impossible.

A causal question asks what would have happened otherwise, and the otherwise is never observed. Every method in this area is a way of constructing a credible counterfactual, and every failure is a counterfactual that was not credible.

Definition 3.1

The treatment effect

Average treatment effect, τ=E[Y(1)−Y(0)]\tau = \mathbb{E}[Y(1) - Y(0)]τ=E[Y(1)−Y(0)] — The difference between the outcome with treatment and without, for the same unit. Only one is ever observed, which is the fundamental problem of causal inference: the data are missing by construction, not by accident.

Proposition 3.2

Selection bias is the default

Whoever received the treatment usually differs from whoever did not, in ways that also affect the outcome. Customers who used the support chat had a problem; patients who took the drug were sicker; funds that report performance are the ones that survived. The naive comparison mixes the treatment effect with those differences and can easily reverse its sign.

Holds when

  • Ask who *chose* to be treated and why — the selection mechanism is the confounder.
  • Controlling for observables handles only the confounders you thought of and measured.
  • When selection is on something unobservable, no regression will fix it and you need a design.
Mild cases: treated93Mild cases: control87Severe cases: treated73Severe cases: control69Everyone: treated78Everyone: control83
Figure 3.3 · A treatment that wins twice and loses overall. The treatment wins in both subgroups and loses in the pool, because it was given to far more of the severe cases. Nothing here is a statistical artefact to be corrected — the grouping variable is the finding, and reporting the pooled number alone would be the error.

Simpson’s paradox is not a paradox. A treatment can help in every subgroup and appear harmful overall, if the subgroups have different base rates and different treatment shares. Nothing contradictory has happened: the aggregate comparison is between different mixtures of subgroups, so it is not answering the question it appears to answer. The lesson is that the correct level of aggregation is determined by the causal structure — condition on what caused the assignment, not on everything available — and that an aggregate difference with no causal story behind it is not evidence of anything.

Example 3.4

Treatment A succeeds in 72% of 100 cases and B in 78% of 100. Split by severity, A wins in both. How?

Show the worked solutionHide the worked solution

Worked solution

  1. Formula
    aggregate=weighted average over subgroups\text{aggregate} = \text{weighted average over subgroups}aggregate=weighted average over subgroups
  2. Substitute
    A: 90/100 severe,B: 10/100 severe\text{A: } 90/100 \text{ severe}, \quad \text{B: } 10/100 \text{ severe}A: 90/100 severe,B: 10/100 severe
  3. Solve
    Severe: A 63/90=70%,B 6/10=60%\text{Severe: A } 63/90 = 70\%, \quad \text{B } 6/10 = 60\%Severe: A 63/90=70%,B 6/10=60%
  4. Mild: A 9/10=90%,B 72/90=80%\text{Mild: A } 9/10 = 90\%, \quad \text{B } 72/90 = 80\%Mild: A 9/10=90%,B 72/90=80%
  5. A overall 72/100,B overall 78/100\text{A overall } 72/100, \quad \text{B overall } 78/100A overall 72/100,B overall 78/100
  6. Answer
    A is better in each group and worse overall\text{A is better in each group and worse overall}A is better in each group and worse overall

Sanity check. A was given overwhelmingly to severe cases and B to mild ones, so the aggregate compares a hard sample against an easy one. If severity determined which treatment was given, the subgroup comparison is the causal one and the aggregate is meaningless.

The rest of this lesson is in Premium

You have read the opening. 11 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Prediction cases: demand, pricing and the traps in eachMarket-data cases: short-horizon prediction from the order book →
On this page
  • The treatment effect
  • Selection bias is the default
  • A treatment that wins twice and loses overall
  • Worked example

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.