The research pipeline: from hypothesis to live capital
SIG · Chapter 112 min readAsked at Two Sigma, AQR, Citadel, Point72
After this lesson you should be able to
- Describe the stages a signal passes through and what kills it at each.
- Say why the hypothesis has to come before the data.
- Explain what research hygiene means in practice.
Signal research is a pipeline with a brutal attrition rate: most ideas die, and the discipline is in killing them cheaply and in the right order. Working the stages in sequence — hypothesis, data, feature, evaluation, risk, capacity, live — means the expensive checks only ever run on ideas that have survived the cheap ones.
| Stage | Question | What kills it here |
|---|---|---|
| Hypothesis | Why should this predict anything? | No economic mechanism |
| Data | Is it available point-in-time? | Look-ahead, survivorship, revisions |
| Feature | How do I express it numerically? | The construction destroys the signal |
| Evaluation | Does it forecast, out of sample? | IC indistinguishable from zero |
| Risk | Is it a known factor in disguise? | Alpha vanishes after neutralisation |
| Capacity | How much can it carry? | Costs exceed the edge at size |
| Live | Does it work with real money? | Live performance well below backtest |
Proposition 1.2
The hypothesis comes first
Write down what you expect to find and why, before touching the data. It is not a formality: it fixes the number of hypotheses being tested at one, which is the only thing that makes a p-value or a Sharpe ratio interpretable. An idea that arrives from a data sweep has an unknown trial count attached, and no amount of subsequent validation recovers it.
Holds when
- A mechanism can be behavioural, structural or a risk premium — but it must exist and be stateable.
- "The data say so" is not a hypothesis; it is a description of a search.
- Pre-registering the universe, the horizon and the evaluation metric is part of the same discipline.
Why the mechanism does the work. There are far more plausible-looking patterns in financial data than there are real ones, so statistics alone cannot separate them — the multiple-testing arithmetic makes that clear. What the mechanism buys is a *prior*: an idea with a reason to work should be believed on much weaker evidence than one without. It also tells you when to abandon it, because a signal with a stated cause has a stated condition under which the cause stops operating. A signal with no mechanism has no such condition, so you learn nothing from its decay.
Research hygiene
- Version the code and the data together, so a result can be reproduced exactly.
- Log every variant tried, including the abandoned ones, and keep the count.
- Build the evaluation before the model, so you cannot tune the evaluation to the result.
- Keep a final holdout untouched until the decision is made.
- Write the conclusion you expect before you run it; note when you were wrong.
Example 1.3
Attrition arithmetic
A team tests 200 ideas a year, and one in twenty is genuinely real. At significance and power, how many findings are false?
Show the worked solutionHide the worked solution
Worked solution
- Formula
- Substitute
- Solve
- Answer
Sanity check. Even with a respectable hit rate and standard thresholds, the majority of significant results are wrong. Raising the prior — by only testing ideas with a mechanism — is far more effective than tightening the threshold, because it changes rather than .
The rest of this lesson is in Premium
You have read the opening. 11 more sections follow, including 3 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.