Skip to content
QuantMax
QuantMax
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • STATStatistics and inference
    • REGRegression and econometrics
    • TSTime series
    • LALinear algebra
    • SCStochastic calculus
    • MLMachine learning
      • 1Framework

        • The framework: bias, variance, capacity and dimensionality
      • 2Tree-based methods

        • Trees: bagging, random forests and gradient boosting
      • 3Other supervised methods

        • Other supervised methods: kNN, SVMs and the kernel trick
      • 4Unsupervised learning

        • Unsupervised learning: clustering assets and correlation structure
      • 5Model selection and evaluation

        • Why k-fold cross-validation is wrong on financial data
      • 6Optimisation for learning

        • Optimisation: gradient descent, momentum and Adam
      • 7Neural networks

        • Neural networks: backpropagation, and when they are the wrong tool
    • SIGAlpha and signal research
    • CASEResearch case studies

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative research
  3. /Machine learning
  4. /Framework

The framework: bias, variance, capacity and dimensionality

ML · Chapter 1·12 min read·Asked at Two Sigma, Citadel, QuantCo, DE Shaw

After this lesson you should be able to

  • Decompose prediction error and say which term dominates on financial data.
  • Relate capacity to overfitting and to the amount of data available.
  • Explain the curse of dimensionality in concrete terms.

Every supervised learning problem is the same trade: a more flexible model fits the training data better and generalises worse. What makes financial data distinctive is where the optimum sits — so far toward the simple end that the usual machine-learning instincts are actively misleading.

Equation 1.1

The error decomposition

Three terms, of which only the first two are under your control.

E[(y−f^(x))2]=Bias2⏟too rigid+Var⏟too flexible+σ2⏟irreducible\mathbb{E}\big[(y - \hat{f}(x))^2\big] = \underbrace{\mathrm{Bias}^2}_{\text{too rigid}} + \underbrace{\mathrm{Var}}_{\text{too flexible}} + \underbrace{\sigma^2}_{\text{irreducible}}E[(y−f^​(x))2]=too rigidBias2​​+too flexibleVar​​+irreducibleσ2​​
σ2\sigma^2σ2
Noise in the target. On daily returns this is around 99%99\%99% of the total.
Var\mathrm{Var}Var
How much the fitted model changes if you resample the training data.

When the noise term dominates everything. The usual picture has bias and variance roughly comparable, with a clear interior optimum. On daily returns the irreducible term is overwhelming: even a perfect model would explain about one per cent of the variance. In that regime almost any added capacity buys a negligible reduction in bias and a large increase in variance, so the optimum sits at a model far simpler than intuition — or than experience with images and text — suggests. "Try a bigger model" is good advice in most of machine learning and bad advice here.

KnobMore capacityLess
Model classDeep network, boosted treesLinear model
Tree depthDeeperStumps
RegularisationSmall λ\lambdaλLarge λ\lambdaλ
FeaturesMore of themA selected few
Training timeMore iterationsEarly stopping
Table 1.2 · What controls capacity. These are substitutes, which is why tuning them jointly by grid search is both expensive and an easy way to overfit the validation set.

Proposition 1.3

The curse of dimensionality

In high dimensions, data becomes sparse and every point becomes roughly equidistant from every other. To cover the unit cube at a given density you need a number of points exponential in the dimension — ten per axis is 10 points in one dimension and 102010^{20}1020 in twenty.

Holds when

  • Distance-based methods like kNN degrade fastest, since "nearest" stops meaning anything.
  • Most real high-dimensional data lies near a lower-dimensional manifold, which is why anything works at all.
  • Adding a useless feature is not free: it adds variance and dilutes every neighbourhood.
0.436012Bias²VarianceTotalModel capacityExpected error
Figure 1.4 · The only curve in machine learning. The minimum is where the two curves cross, not where either is small. Everything that calls itself regularisation — fewer trees, a shorter depth, a larger λ\lambdaλ, early stopping — is a way of sliding left along this axis, and the right amount is an empirical question every time.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read The law of large numbers and the central limit theorem in full, free.

Trees: bagging, random forests and gradient boosting →
On this page
  • The error decomposition
  • What controls capacity
  • The curse of dimensionality
  • The only curve in machine learning

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.