Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
      • 1Limit theorems

        • The law of large numbers and the central limit theorem
      • 2Estimation

        • Estimation: bias, variance and maximum likelihood
      • 3Confidence intervals

        • Confidence intervals, and the interval you would trade
      • 4Hypothesis testing

        • Hypothesis testing: errors, power and which test to use
      • 5p-values and multiple testing

        • p-values, p-hacking and the multiple-testing problem
      • 6Resampling

        • Resampling: the bootstrap, permutation tests and where they break
      • 7Bayesian statistics

        • Bayesian statistics: conjugacy, shrinkage and credible intervals
      • 8Experiment design

        • Experiment design: randomisation, peeking and minimum detectable effect
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Statistics and inference
  4. /Experiment design

Experiment design: randomisation, peeking and minimum detectable effect

STAT · Chapter 8·12 min read·Asked at QuantCo, Two Sigma, Citadel, Point72

Assumes Hypothesis testing: errors, power and which test to use.

After this lesson you should be able to

  • Design an experiment from a minimum detectable effect backwards.
  • Say what randomisation buys and what it does not.
  • Handle sequential testing without inflating the error rate.

An experiment is the only clean way to establish a causal effect, and almost all of its value is decided before any data arrive. Randomisation buys comparability, the sample size buys sensitivity, and a pre-registered analysis plan buys the right to believe the result.

Proposition 8.1

What randomisation buys

Random assignment makes the treatment independent of everything else — observed and unobserved. That is why an experiment answers a causal question that no amount of regression on observational data can: you are not controlling for confounders, you are making them irrelevant in expectation.

Holds when

  • It balances confounders *in expectation*, not in any particular sample — check balance afterwards.
  • Stratify or block on variables you know matter, which removes their variance rather than hoping it balances.
  • Randomise at the level at which interference happens: if users influence each other, randomise by market or by cluster.

Equation 8.2

Minimum detectable effect

Run the power calculation the other way: given the sample you can actually get, what is the smallest effect you could reliably detect?

MDE=(zα/2+zβ)2σ2n\mathrm{MDE} = (z_{\alpha/2} + z_\beta)\sqrt{\frac{2\sigma^2}{n}}MDE=(zα/2​+zβ​)n2σ2​​
nnn
Per arm.
zα/2+zβ=2.80z_{\alpha/2} + z_\beta = 2.80zα/2​+zβ​=2.80
The usual 5%5\%5% and 80%80\%80% design.

Compute the MDE before you run anything. The most useful number in experiment design is the one that tells you not to bother. If the traffic you have supports detecting only a ten per cent lift and the intervention plausibly delivers one, the experiment cannot succeed — it will return "not significant" regardless of whether the effect is real, and that outcome carries no information. Computing the MDE first converts an argument about results into a decision about whether the experiment is worth running, which is a far better conversation to have.

0.010.020.05050000100000n ∝ 1/δ²Minimum detectable effectSample size per arm
Figure 8.3 · What halving the effect you can detect costs. At a 20%20\%20% base rate, detecting two percentage points needs about 6,3006{,}3006,300 per arm and detecting one point needs four times that. Sample size scales with the inverse square of the effect, which is why the effect you care about has to be chosen before the experiment rather than discovered after it.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Bayesian statistics: conjugacy, shrinkage and credible intervalsBack to Statistics and inference →
On this page
  • What randomisation buys
  • Minimum detectable effect
  • What halving the effect you can detect costs

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.