Causal cases: treatment effects, confounding and surprising results
CASE · Chapter 312 min readAsked at QuantCo, Two Sigma, Citadel, Point72
Assumes Prediction cases: demand, pricing and the traps in each.
After this lesson you should be able to
- State what a treatment effect is and why it needs a counterfactual.
- Recognise selection bias, Simpson’s paradox and survivorship in a case.
- Say what you would do when randomisation is impossible.
A causal question asks what would have happened otherwise, and the otherwise is never observed. Every method in this area is a way of constructing a credible counterfactual, and every failure is a counterfactual that was not credible.
Definition 3.1
The treatment effect
Average treatment effect, — The difference between the outcome with treatment and without, for the same unit. Only one is ever observed, which is the fundamental problem of causal inference: the data are missing by construction, not by accident.
Proposition 3.2
Selection bias is the default
Whoever received the treatment usually differs from whoever did not, in ways that also affect the outcome. Customers who used the support chat had a problem; patients who took the drug were sicker; funds that report performance are the ones that survived. The naive comparison mixes the treatment effect with those differences and can easily reverse its sign.
Holds when
- Ask who *chose* to be treated and why — the selection mechanism is the confounder.
- Controlling for observables handles only the confounders you thought of and measured.
- When selection is on something unobservable, no regression will fix it and you need a design.
Simpson’s paradox is not a paradox. A treatment can help in every subgroup and appear harmful overall, if the subgroups have different base rates and different treatment shares. Nothing contradictory has happened: the aggregate comparison is between different mixtures of subgroups, so it is not answering the question it appears to answer. The lesson is that the correct level of aggregation is determined by the causal structure — condition on what caused the assignment, not on everything available — and that an aggregate difference with no causal story behind it is not evidence of anything.
Example 3.4
Treatment A succeeds in 72% of 100 cases and B in 78% of 100. Split by severity, A wins in both. How?
Show the worked solutionHide the worked solution
Worked solution
- Formula
- Substitute
- Solve
- Answer
Sanity check. A was given overwhelmingly to severe cases and B to mild ones, so the aggregate compares a hard sample against an easy one. If severity determined which treatment was given, the subgroup comparison is the causal one and the aggregate is meaningless.
The rest of this lesson is in Premium
You have read the opening. 11 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.