Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
      • 1Limit theorems

        • The law of large numbers and the central limit theorem
      • 2Estimation

        • Estimation: bias, variance and maximum likelihood
      • 3Confidence intervals

        • Confidence intervals, and the interval you would trade
      • 4Hypothesis testing

        • Hypothesis testing: errors, power and which test to use
      • 5p-values and multiple testing

        • p-values, p-hacking and the multiple-testing problem
      • 6Resampling

        • Resampling: the bootstrap, permutation tests and where they break
      • 7Bayesian statistics

        • Bayesian statistics: conjugacy, shrinkage and credible intervals
      • 8Experiment design

        • Experiment design: randomisation, peeking and minimum detectable effect
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Statistics and inference
  4. /Hypothesis testing

Hypothesis testing: errors, power and which test to use

STAT · Chapter 4·12 min read·Asked at Two Sigma, Citadel, QuantCo, DE Shaw

Assumes Confidence intervals, and the interval you would trade.

After this lesson you should be able to

  • Distinguish the two error types and say which costs more in context.
  • Compute the power of a test and the sample size it needs.
  • Choose the right test for a given comparison.

A hypothesis test is a decision rule with two ways to be wrong, and the whole design consists of choosing how much of each you will tolerate. In research the type that matters is usually the one nobody controls: failing to detect a real effect because the sample was never large enough.

DecisionNull is trueNull is false
RejectType I error, probability α\alphaαCorrect — power, 1−β1 - \beta1−β
Fail to rejectCorrectType II error, probability β\betaβ
Table 4.1 · The two errors. You choose α\alphaα directly. β\betaβ follows from the effect size, the variance and the sample size — which is why power is something you design for rather than observe.

Equation 4.2

Sample size from power

The two-sample size needed to detect a difference δ\deltaδ with power 1−β1-\beta1−β at level α\alphaα.

n≈2(zα/2+zβ)2σ2δ2n \approx \frac{2\left(z_{\alpha/2} + z_{\beta}\right)^2\sigma^2}{\delta^2}n≈δ22(zα/2​+zβ​)2σ2​
zα/2+zβz_{\alpha/2} + z_\betazα/2​+zβ​
For α=5%\alpha = 5\%α=5% and 80%80\%80% power this is 1.96+0.84=2.801.96 + 0.84 = 2.801.96+0.84=2.80.
δ\deltaδ
The minimum effect you care about detecting, chosen before you look at the data.

Example 4.3

You want to detect a 2%2\%2% improvement in a conversion rate from a base of 20%20\%20%, at 5%5\%5% significance and 80%80\%80% power. How many per arm?

Show the worked solutionHide the worked solution

Worked solution

  1. Formula
    n=2(zα/2+zβ)2p(1−p)δ2n = \frac{2(z_{\alpha/2} + z_\beta)^2 p(1-p)}{\delta^2}n=δ22(zα/2​+zβ​)2p(1−p)​
  2. Substitute
    p=0.20, δ=0.02, (1.96+0.84)2=7.84p = 0.20, \ \delta = 0.02, \ (1.96 + 0.84)^2 = 7.84p=0.20, δ=0.02, (1.96+0.84)2=7.84
  3. Solve
    p(1−p)=0.16p(1-p) = 0.16p(1−p)=0.16
  4. n=2×7.84×0.160.0004n = \frac{2 \times 7.84 \times 0.16}{0.0004}n=0.00042×7.84×0.16​
  5. Answer
    ≈6,300 per arm\approx 6{,}300 \text{ per arm}≈6,300 per arm

Sanity check. Halving the effect you want to detect quadruples the sample, since δ\deltaδ enters squared. That single fact explains why most business experiments are underpowered: the effects are small and nobody budgets for the square.

-3-2-10123−1.961.96The test statistic under the null
Figure 4.4 · What a 5%5\%5% test actually rejects. Both shaded tails together hold five per cent of the distribution, which is what α=5%\alpha = 5\%α=5% buys you: reject outside ±1.96\pm 1.96±1.96 and you will be wrong one time in twenty when the null is true. Power is a separate picture — it asks how much of a *different* distribution lands in the same two tails.
QuestionTestAssumes
Mean against a value, σ\sigmaσ knownzzz-testNormal or large nnn
Mean against a value, σ\sigmaσ estimatedttt-testNormal or large nnn
Two meansTwo-sample tttIndependent groups; Welch if variances differ
Before and after on the same unitsPaired tttDifferences are what you test
Two proportionsTwo-proportion zzzEnough successes in each arm
VariancesFFF-testVery sensitive to non-normality
Several meansANOVAEqual variances; follow up carefully
Categorical associationChi-squareExpected counts above about five
No distributional assumptionPermutation testExchangeability under the null
Table 4.5 · Which test. The paired row is the one candidates miss. When the same units are measured twice, pairing removes the between-unit variance and can cut the required sample by an order of magnitude.

The rest of this lesson is in Premium

You have read the opening. 9 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Confidence intervals, and the interval you would tradep-values, p-hacking and the multiple-testing problem →
On this page
  • The two errors
  • Sample size from power
  • Worked example
  • What a 5%5\%5% test actually rejects
  • Which test

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.