The framework: bias, variance, capacity and dimensionality
ML · Chapter 112 min readAsked at Two Sigma, Citadel, QuantCo, DE Shaw
After this lesson you should be able to
- Decompose prediction error and say which term dominates on financial data.
- Relate capacity to overfitting and to the amount of data available.
- Explain the curse of dimensionality in concrete terms.
Every supervised learning problem is the same trade: a more flexible model fits the training data better and generalises worse. What makes financial data distinctive is where the optimum sits — so far toward the simple end that the usual machine-learning instincts are actively misleading.
Equation 1.1
The error decomposition
Three terms, of which only the first two are under your control.
- Noise in the target. On daily returns this is around of the total.
- How much the fitted model changes if you resample the training data.
When the noise term dominates everything. The usual picture has bias and variance roughly comparable, with a clear interior optimum. On daily returns the irreducible term is overwhelming: even a perfect model would explain about one per cent of the variance. In that regime almost any added capacity buys a negligible reduction in bias and a large increase in variance, so the optimum sits at a model far simpler than intuition — or than experience with images and text — suggests. "Try a bigger model" is good advice in most of machine learning and bad advice here.
| Knob | More capacity | Less |
|---|---|---|
| Model class | Deep network, boosted trees | Linear model |
| Tree depth | Deeper | Stumps |
| Regularisation | Small | Large |
| Features | More of them | A selected few |
| Training time | More iterations | Early stopping |
Proposition 1.3
The curse of dimensionality
In high dimensions, data becomes sparse and every point becomes roughly equidistant from every other. To cover the unit cube at a given density you need a number of points exponential in the dimension — ten per axis is 10 points in one dimension and in twenty.
Holds when
- Distance-based methods like kNN degrade fastest, since "nearest" stops meaning anything.
- Most real high-dimensional data lies near a lower-dimensional manifold, which is why anything works at all.
- Adding a useless feature is not free: it adds variance and dilutes every neighbourhood.
The rest of this lesson is in Premium
You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.