Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
      • 1Framework

        • The framework: bias, variance, capacity and dimensionality
      • 2Tree-based methods

        • Trees: bagging, random forests and gradient boosting
      • 3Other supervised methods

        • Other supervised methods: kNN, SVMs and the kernel trick
      • 4Unsupervised learning

        • Unsupervised learning: clustering assets and correlation structure
      • 5Model selection and evaluation

        • Why k-fold cross-validation is wrong on financial data
      • 6Optimisation for learning

        • Optimisation: gradient descent, momentum and Adam
      • 7Neural networks

        • Neural networks: backpropagation, and when they are the wrong tool
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Machine learning
  4. /Neural networks

Neural networks: backpropagation, and when they are the wrong tool

ML · Chapter 7·12 min read·Asked at Two Sigma, Citadel, QuantCo, Jump

Assumes Optimisation: gradient descent, momentum and Adam.

After this lesson you should be able to

  • Explain backpropagation as the chain rule applied efficiently.
  • Name the regularisers and what each one does.
  • Say where deep learning wins in finance and where it does not.

A neural network is a composition of linear maps and non-linearities, fitted by gradient descent. The mathematics is the chain rule organised carefully; the interesting judgement is whether the problem has the compositional structure that makes the architecture worth its enormous capacity.

Equation 7.1

A feedforward layer

A linear map followed by an element-wise non-linearity. Stack them and you have a network; remove the non-linearity and the whole stack collapses to a single linear map.

a(l)=ϕ ⁣(W(l)a(l−1)+b(l))a^{(l)} = \phi\!\left(W^{(l)}a^{(l-1)} + b^{(l)}\right)a(l)=ϕ(W(l)a(l−1)+b(l))
ϕ\phiϕ
Activation — ReLU by default, since it does not saturate and its gradient is trivial.
W(l)W^{(l)}W(l)
Weights of layer lll, learned by gradient descent.

Derivation 7.2

Backpropagation

Compute every partial derivative in one backward sweep rather than one forward sweep per parameter.

  1. δ(L)=∇aL⊙ϕ′(z(L))\delta^{(L)} = \nabla_a L \odot \phi'(z^{(L)})δ(L)=∇a​L⊙ϕ′(z(L))

    Start at the output with the loss gradient.

  2. δ(l)=(W(l+1)⊤δ(l+1))⊙ϕ′(z(l))\delta^{(l)} = \left(W^{(l+1)\top}\delta^{(l+1)}\right) \odot \phi'(z^{(l)})δ(l)=(W(l+1)⊤δ(l+1))⊙ϕ′(z(l))

    Propagate the error backwards through the transpose of each weight matrix.

  3. ∂L∂W(l)=δ(l)a(l−1)⊤\frac{\partial L}{\partial W^{(l)}} = \delta^{(l)}a^{(l-1)\top}∂W(l)∂L​=δ(l)a(l−1)⊤
all gradients at roughly the cost of one forward pass\text{all gradients at roughly the cost of one forward pass}all gradients at roughly the cost of one forward pass
Backpropagation2One finite difference per weight2,000,000
Figure 7.3 · Why nobody differentiates a network by hand. A million-parameter model, one gradient. Backpropagation reuses the forward pass and walks the chain rule backwards once; finite differences re-evaluate the whole network for every weight. The same trick, under the name adjoint mode, is how a bank gets all its Greeks in one sweep.

Why backpropagation is the whole trick. A network may have millions of parameters, and computing each derivative by perturbing it individually would cost a forward pass each — completely infeasible. Backpropagation exploits the fact that the derivatives share almost all of their work: the chain rule from the loss to any weight passes through the same intermediate quantities, so computing them once and reusing them gives every gradient in one sweep. It is dynamic programming on the computation graph, and it is what made deep learning possible at all.

TechniqueWhat it doesWhy it helps
Weight decayL2L_2L2​ penalty on the weightsSame as ridge — shrinks toward zero
DropoutRandomly zero units during trainingForces redundancy; an implicit ensemble
Early stoppingHalt when validation stops improvingLimits effective capacity
Batch normalisationStandardise activations per batchConditions the optimisation; mildly regularises
Data augmentationGenerate plausible variantsRarely available for financial data
Table 7.4 · The regularisers. The last row is the quiet reason deep learning transfers badly to finance: images can be rotated and cropped to manufacture more data, and a return series cannot.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 5 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

← Optimisation: gradient descent, momentum and AdamBack to Machine learning →
On this page
  • A feedforward layer
  • Backpropagation
  • Why nobody differentiates a network by hand
  • The regularisers

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.