Trees: bagging, random forests and gradient boosting
ML · Chapter 212 min readAsked at Two Sigma, Citadel, QuantCo, AQR
Assumes The framework: bias, variance, capacity and dimensionality.
After this lesson you should be able to
- Say what a tree splits on and why a single tree overfits.
- Distinguish bagging from boosting by which error term they attack.
- Read a feature importance sceptically.
A decision tree partitions the feature space into boxes and predicts a constant in each. Alone it is high variance and nearly useless; combined by bagging or boosting it becomes the strongest general-purpose method on tabular data, which is most of what quantitative research works with.
Proposition 2.1
How a tree is grown
At each node, search every feature and every threshold for the split that most reduces impurity — variance for regression, Gini or entropy for classification — then recurse. Grown without limit a tree fits the training data exactly, which is why depth, minimum leaf size or pruning is always applied.
Holds when
- Splits are axis-aligned, so a diagonal boundary needs a staircase of them.
- Trees are invariant to monotone transformations of a feature, so scaling and log transforms change nothing.
- Interactions come free: a split below another split is a conditional effect.
| Aspect | Bagging / random forest | Gradient boosting |
|---|---|---|
| Attacks | Variance | Bias |
| Base learners | Deep trees, fitted independently | Shallow trees, fitted sequentially |
| Each tree sees | A bootstrap sample and a feature subset | The residuals of everything so far |
| Parallelisable | Yes | No — inherently sequential |
| Overfits by | Barely, with more trees | Readily, with too many rounds |
| Main knobs | Number of trees, features per split | Learning rate, depth, number of rounds |
Why averaging decorrelated trees works. Averaging independent estimates cuts the variance by ; averaging perfectly correlated ones cuts it by nothing. Bagged trees fitted on bootstrap samples are already somewhat decorrelated, but they still tend to split on the same dominant feature first. Random forests add a second randomisation — considering only a subset of features at each split — precisely to force the trees apart. The entire innovation is decorrelation, and it is the same diversification argument as a portfolio.
The rest of this lesson is in Premium
You have read the opening. 11 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.