Experiment design: randomisation, peeking and minimum detectable effect
STAT · Chapter 812 min readAsked at QuantCo, Two Sigma, Citadel, Point72
Assumes Hypothesis testing: errors, power and which test to use.
After this lesson you should be able to
- Design an experiment from a minimum detectable effect backwards.
- Say what randomisation buys and what it does not.
- Handle sequential testing without inflating the error rate.
An experiment is the only clean way to establish a causal effect, and almost all of its value is decided before any data arrive. Randomisation buys comparability, the sample size buys sensitivity, and a pre-registered analysis plan buys the right to believe the result.
Proposition 8.1
What randomisation buys
Random assignment makes the treatment independent of everything else — observed and unobserved. That is why an experiment answers a causal question that no amount of regression on observational data can: you are not controlling for confounders, you are making them irrelevant in expectation.
Holds when
- It balances confounders *in expectation*, not in any particular sample — check balance afterwards.
- Stratify or block on variables you know matter, which removes their variance rather than hoping it balances.
- Randomise at the level at which interference happens: if users influence each other, randomise by market or by cluster.
Equation 8.2
Minimum detectable effect
Run the power calculation the other way: given the sample you can actually get, what is the smallest effect you could reliably detect?
- Per arm.
- The usual and design.
Compute the MDE before you run anything. The most useful number in experiment design is the one that tells you not to bother. If the traffic you have supports detecting only a ten per cent lift and the intervention plausibly delivers one, the experiment cannot succeed — it will return "not significant" regardless of whether the effect is real, and that outcome carries no information. Computing the MDE first converts an argument about results into a decision about whether the experiment is worth running, which is a far better conversation to have.
The rest of this lesson is in Premium
You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.