Hypothesis testing: errors, power and which test to use
STAT · Chapter 412 min readAsked at Two Sigma, Citadel, QuantCo, DE Shaw
Assumes Confidence intervals, and the interval you would trade.
After this lesson you should be able to
- Distinguish the two error types and say which costs more in context.
- Compute the power of a test and the sample size it needs.
- Choose the right test for a given comparison.
A hypothesis test is a decision rule with two ways to be wrong, and the whole design consists of choosing how much of each you will tolerate. In research the type that matters is usually the one nobody controls: failing to detect a real effect because the sample was never large enough.
| Decision | Null is true | Null is false |
|---|---|---|
| Reject | Type I error, probability | Correct — power, |
| Fail to reject | Correct | Type II error, probability |
Equation 4.2
Sample size from power
The two-sample size needed to detect a difference with power at level .
- For and power this is .
- The minimum effect you care about detecting, chosen before you look at the data.
Example 4.3
You want to detect a improvement in a conversion rate from a base of , at significance and power. How many per arm?
Show the worked solutionHide the worked solution
Worked solution
- Formula
- Substitute
- Solve
- Answer
Sanity check. Halving the effect you want to detect quadruples the sample, since enters squared. That single fact explains why most business experiments are underpowered: the effects are small and nobody budgets for the square.
| Question | Test | Assumes |
|---|---|---|
| Mean against a value, known | -test | Normal or large |
| Mean against a value, estimated | -test | Normal or large |
| Two means | Two-sample | Independent groups; Welch if variances differ |
| Before and after on the same units | Paired | Differences are what you test |
| Two proportions | Two-proportion | Enough successes in each arm |
| Variances | -test | Very sensitive to non-normality |
| Several means | ANOVA | Equal variances; follow up carefully |
| Categorical association | Chi-square | Expected counts above about five |
| No distributional assumption | Permutation test | Exchangeability under the null |
The rest of this lesson is in Premium
You have read the opening. 9 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.