Neural networks: backpropagation, and when they are the wrong tool
ML · Chapter 712 min readAsked at Two Sigma, Citadel, QuantCo, Jump
After this lesson you should be able to
- Explain backpropagation as the chain rule applied efficiently.
- Name the regularisers and what each one does.
- Say where deep learning wins in finance and where it does not.
A neural network is a composition of linear maps and non-linearities, fitted by gradient descent. The mathematics is the chain rule organised carefully; the interesting judgement is whether the problem has the compositional structure that makes the architecture worth its enormous capacity.
Equation 7.1
A feedforward layer
A linear map followed by an element-wise non-linearity. Stack them and you have a network; remove the non-linearity and the whole stack collapses to a single linear map.
- Activation — ReLU by default, since it does not saturate and its gradient is trivial.
- Weights of layer , learned by gradient descent.
Derivation 7.2
Backpropagation
Compute every partial derivative in one backward sweep rather than one forward sweep per parameter.
Start at the output with the loss gradient.
Propagate the error backwards through the transpose of each weight matrix.
Why backpropagation is the whole trick. A network may have millions of parameters, and computing each derivative by perturbing it individually would cost a forward pass each — completely infeasible. Backpropagation exploits the fact that the derivatives share almost all of their work: the chain rule from the loss to any weight passes through the same intermediate quantities, so computing them once and reusing them gives every gradient in one sweep. It is dynamic programming on the computation graph, and it is what made deep learning possible at all.
| Technique | What it does | Why it helps |
|---|---|---|
| Weight decay | penalty on the weights | Same as ridge — shrinks toward zero |
| Dropout | Randomly zero units during training | Forces redundancy; an implicit ensemble |
| Early stopping | Halt when validation stops improving | Limits effective capacity |
| Batch normalisation | Standardise activations per batch | Conditions the optimisation; mildly regularises |
| Data augmentation | Generate plausible variants | Rarely available for financial data |
The rest of this lesson is in Premium
You have read the opening. 10 more sections follow, including 5 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.