p-values, p-hacking and the multiple-testing problem
STAT · Chapter 513 min readAsked at Two Sigma, Citadel, DE Shaw, AQR
Assumes The law of large numbers and the central limit theorem.
After this lesson you should be able to
- State precisely what a p-value is and list three things it is not.
- Compute the family-wise error rate for a set of independent tests.
- Choose between Bonferroni and false-discovery-rate control, and say why.
Signal research is thousands of hypothesis tests wearing a trench coat. Understanding what a p-value measures, and what happens to it when you run a test ten thousand times, is the difference between a research process that discovers things and one that manufactures them.
Definition 5.1
What a p-value is
p-value, — The probability of seeing data at least as extreme as what you observed, *given that the null hypothesis is true*. It is a statement about data under an assumption, not about the assumption.
Proposition 5.2
Three things it is not
A p-value is not the probability the null is true. It is not the probability your result is a fluke. And one minus the p-value is not the probability the effect is real. All three require a prior, which the p-value does not have.
Holds when
- is not — this is the same inversion error as the disease-testing brainteaser.
- A p-value of 0.05 is routinely consistent with a posterior probability of the null well above 20%.
Equation 5.3
The family-wise error rate
The probability of at least one false positive across independent tests at level .
- The number of hypotheses tested — including the ones you tried and discarded.
- The per-test significance level.
Example 5.4
You backtest 100 signals, none of which has any edge, and keep any that is significant at the level. How many do you expect to keep, and what is the chance you keep at least one?
Show the worked solutionHide the worked solution
Worked solution
- Formula
- Substitute
- Solve
- Answer
Sanity check. Every one of those five will have a plausible story attached, because you will construct one. That is the whole problem.
- Clears the bar anyway — 5 of 100
- Correctly discarded — 95 of 100
Two different fears. Bonferroni controls the chance of making even one false discovery; Benjamini–Hochberg controls the expected share of your discoveries that are false. A researcher running thousands of backtests can live with one live signal in twenty being junk, but cannot live with a procedure so strict it never rejects anything. Which fear you are managing decides which correction is right.
The rest of this lesson is in Premium
You have read the opening. 12 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.