Skip to content
QuantMax
QuantMax
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • GAMEGames, decision theory and puzzles
    • MMMarket making
    • MKTMarkets and products

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Formula reference

Machine learning

7 lessons · 10 equations. Each lesson below gives its formulas and key rules; open the lesson for the full explanation.

The framework: bias, variance, capacity and dimensionality

The error decomposition

E[(y−f^(x))2]=Bias2⏟too rigid+Var⏟too flexible+σ2⏟irreducible\mathbb{E}\big[(y - \hat{f}(x))^2\big] = \underbrace{\mathrm{Bias}^2}_{\text{too rigid}} + \underbrace{\mathrm{Var}}_{\text{too flexible}} + \underbrace{\sigma^2}_{\text{irreducible}}E[(y−f^​(x))2]=too rigidBias2​​+too flexibleVar​​+irreducibleσ2​​

Three terms, of which only the first two are under your control.

How much worse a model does out of sample

E[train MSE]=σ2(1−pn),E[test MSE]=σ2(1+pn−p−1)\mathbb{E}[\text{train MSE}] = \sigma^2\left(1 - \frac pn\right), \qquad \mathbb{E}[\text{test MSE}] = \sigma^2\left(1 + \frac{p}{n - p - 1}\right)E[train MSE]=σ2(1−np​),E[test MSE]=σ2(1+n−p−1p​)

For least squares with ppp predictors, nnn observations, noise variance σ2\sigma^2σ2 and Gaussian regressors, when the true coefficients are all zero. Training error flatters the model by the fraction p/np/np/n; test error is inflated by estimation noise in every coefficient.

Remember

  • Error is bias squared plus variance plus irreducible noise.

Trees: bagging, random forests and gradient boosting

Out-of-bag error comes free

Pr⁡(row i not in a bootstrap sample)=(1−1n)n≈e−1≈0.368\Pr(\text{row } i \text{ not in a bootstrap sample}) = \left(1 - \tfrac1n\right)^n \approx e^{-1} \approx 0.368Pr(row i not in a bootstrap sample)=(1−n1​)n≈e−1≈0.368

Each tree in a random forest sees about 63%63\%63% of the rows, so each row can be predicted by the roughly third of trees that never saw it. Averaging those predictions gives an almost unbiased estimate of test error without a separate validation set — though on time series, where nearby rows are correlated, it is optimistic.

Remember

  • Trees split axis-aligned and are invariant to monotone feature transformations.

Other supervised methods: kNN, SVMs and the kernel trick

The soft-margin SVM

min⁡w,b 12∥w∥2+C∑imax⁡(0, 1−yi(w⊤xi+b))\min_{w, b}\ \tfrac{1}{2}\lVert w\rVert^2 + C\sum_i \max\big(0,\ 1 - y_i(w^\top x_i + b)\big)w,bmin​ 21​∥w∥2+Ci∑​max(0, 1−yi​(w⊤xi​+b))

The hinge loss penalises points inside the margin or on the wrong side; CCC trades margin width against violations. Large CCC fits the training data hard; small CCC keeps a wide, stable margin — the SVM’s version of regularisation.

Remember

  • kNN assumes locality and a meaningful distance; standardise, and beware dimension.

Unsupervised learning: clustering assets and correlation structure

Correlation as a distance

dij=2(1−ρij)d_{ij} = \sqrt{2(1 - \rho_{ij})}dij​=2(1−ρij​)​

The standard conversion. It is a proper metric — it satisfies the triangle inequality — which naive alternatives like 1−ρ1 - \rho1−ρ do not.

Remember

  • k-means assumes round, equal-sized Euclidean clusters; correlations are not Euclidean.

Why k-fold cross-validation is wrong on financial data

The standard error of an AUC

SE⁡(A)=A(1−A)+(n1−1)(Q1−A2)+(n2−1)(Q2−A2)n1n2, Q1=A2−A, Q2=2A21+A\operatorname{SE}(A) = \sqrt{\frac{A(1-A) + (n_1 - 1)(Q_1 - A^2) + (n_2 - 1)(Q_2 - A^2)}{n_1n_2}},\ Q_1 = \frac{A}{2 - A},\ Q_2 = \frac{2A^2}{1 + A}SE(A)=n1​n2​A(1−A)+(n1​−1)(Q1​−A2)+(n2​−1)(Q2​−A2)​​, Q1​=2−AA​, Q2​=1+A2A2​

Hanley and McNeil’s approximation, for n1n_1n1​ positives and n2n_2n2​ negatives treated as independent. With overlapping labels the effective counts are smaller and the error larger.

Remember

  • Financial rows are ordered and overlapping, so k-fold leaks.

Optimisation: gradient descent, momentum and Adam

Stochastic gradient descent

θt+1=θt−η ∇θL(θt;Bt)\theta_{t+1} = \theta_t - \eta\,\nabla_\theta L(\theta_t; \mathcal{B}_t)θt+1​=θt​−η∇θ​L(θt​;Bt​)

Step downhill using the gradient computed on a small batch rather than on the whole dataset.

Momentum turns $\kappa$ into $\sqrt\kappa$

GD: (κ−1κ+1)t,heavy ball: (κ−1κ+1)t\text{GD: } \left(\frac{\kappa - 1}{\kappa + 1}\right)^t, \qquad \text{heavy ball: } \left(\frac{\sqrt\kappa - 1}{\sqrt\kappa + 1}\right)^tGD: (κ+1κ−1​)t,heavy ball: (κ​+1κ​−1​)t

On a quadratic with condition number κ\kappaκ, the best fixed step of gradient descent shrinks the error by the first factor per step; optimally tuned momentum by the second. Replacing κ\kappaκ by its square root is the whole value of momentum on ill-conditioned problems.

Remember

  • Stochastic gradients are noisier per step and far more efficient per unit of computation.

Neural networks: backpropagation, and when they are the wrong tool

A feedforward layer

a(l)=ϕ ⁣(W(l)a(l−1)+b(l))a^{(l)} = \phi\!\left(W^{(l)}a^{(l-1)} + b^{(l)}\right)a(l)=ϕ(W(l)a(l−1)+b(l))

A linear map followed by an element-wise non-linearity. Stack them and you have a network; remove the non-linearity and the whole stack collapses to a single linear map.

The gradient of softmax with cross-entropy

∂∂z(−∑kykln⁡softmax⁡(z)k)=softmax⁡(z)−y\frac{\partial}{\partial z}\Big(-\sum_k y_k\ln\operatorname{softmax}(z)_k\Big) = \operatorname{softmax}(z) - y∂z∂​(−k∑​yk​lnsoftmax(z)k​)=softmax(z)−y

Predicted probabilities minus the one-hot target: the cleanest gradient in deep learning, and the reason softmax and cross-entropy are always paired. Each logit is pushed down in proportion to the probability wrongly assigned to it.

Remember

  • A network without non-linearities collapses to a linear map.

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.