Resampling: the bootstrap, permutation tests and where they break
STAT · Chapter 612 min readAsked at Two Sigma, QuantCo, AQR, Citadel
Assumes Hypothesis testing: errors, power and which test to use.
After this lesson you should be able to
- Describe the bootstrap and what it estimates.
- Use a permutation test where no distributional assumption is available.
- Say why the ordinary bootstrap fails on a time series, and what replaces it.
Resampling replaces a distributional assumption with computation: instead of deriving the sampling distribution of a statistic, you generate it by re-drawing from the data you have. It is the right tool when the statistic is awkward — a median, a Sharpe ratio, a maximum drawdown — and the wrong one the moment the observations are not exchangeable.
Definition 6.1
The bootstrap
Bootstrap, — Treat the sample as if it were the population, draw new samples of the same size from it with replacement, and compute your statistic on each. The spread of those values estimates the sampling distribution — which gives you a standard error or a confidence interval for any statistic at all, including ones with no closed form.
Why re-using the same data is not circular. The obvious objection is that you cannot manufacture information by shuffling what you already have, and you cannot. What the bootstrap estimates is not the parameter — that still comes from the sample — but the *variability* of the estimate, and for that the empirical distribution is a legitimate stand-in for the true one. The plug-in step is the whole trick: the relationship between population and sample is approximated by the relationship between sample and resample.
| Method | Question it answers | Note |
|---|---|---|
| Bootstrap | How variable is my estimate? | Works for almost any statistic |
| Jackknife | Same, by leaving one out at a time | Cheaper, and fails on non-smooth statistics like the median |
| Permutation test | Could this difference be chance? | Exact under exchangeability; no distribution assumed |
| Block bootstrap | Variability with dependent data | Resample blocks, not points |
| Stationary bootstrap | Same, with random block lengths | Avoids artefacts at fixed block boundaries |
Derivation 6.4
A permutation test
Under the null the labels carry no information, so any relabelling is equally likely.
Compute the statistic on the real labelling.
This generates the null distribution directly.
The s include the observed value and keep the test valid.
The rest of this lesson is in Premium
You have read the opening. 11 more sections follow, including 5 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.