Skip to content
QuantMax
QuantMax
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • GAMEGames, decision theory and puzzles
    • MMMarket making
    • MKTMarkets and products

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Formula reference

Time series

7 lessons · 15 equations. Each lesson below gives its formulas and key rules; open the lesson for the full explanation.

Foundations: stationarity, autocorrelation and the Wold decomposition

Standard errors for autocorrelations beyond white noise

Var⁡(ρ^h)≈1T(1+2∑k=1qρk2),h>q\operatorname{Var}(\hat\rho_h) \approx \frac{1}{T}\Big(1 + 2\sum_{k=1}^{q}\rho_k^2\Big),\quad h > qVar(ρ^​h​)≈T1​(1+2k=1∑q​ρk2​),h>q

If the true process is MA(qqq), sample autocorrelations beyond lag qqq are noisier than 1/T1/\sqrt T1/T​ suggests, because earlier dependence propagates into their estimates. Using the white-noise band for every lag finds spurious structure in persistent series.

Remember

  • Weak stationarity: constant mean and variance, covariance depending only on the lag.

The ARMA family: identification, estimation and forecasting

The models

xt=∑i=1pϕixt−i+εt+∑j=1qθjεt−jx_t = \sum_{i=1}^{p}\phi_i x_{t-i} + \varepsilon_t + \sum_{j=1}^{q}\theta_j\varepsilon_{t-j}xt​=i=1∑p​ϕi​xt−i​+εt​+j=1∑q​θj​εt−j​

AR terms carry the series’ own memory; MA terms carry the memory of past shocks.

How much of the variance a forecast leaves

Var⁡(et+h)Var⁡(xt)=1−ϕ2h  (AR(1))\frac{\operatorname{Var}(e_{t+h})}{\operatorname{Var}(x_t)} = 1 - \phi^{2h}\ \ \text{(AR(1))}Var(xt​)Var(et+h​)​=1−ϕ2h  (AR(1))

The fraction of the unconditional variance that remains unexplained hhh steps ahead. At h=1h = 1h=1 it is 1−ϕ21 - \phi^21−ϕ2; as hhh grows it approaches one, and the forecast becomes the mean.

Persistence is underestimated in short samples

E[ϕ^]≈ϕ−1+3ϕT\mathbb{E}[\hat\phi] \approx \phi - \frac{1 + 3\phi}{T}E[ϕ^​]≈ϕ−T1+3ϕ​

The least-squares estimate of an AR(1) coefficient, with an estimated mean, is biased toward zero by roughly this amount (Kendall; Marriott and Pope). The bias is largest precisely for persistent series.

Remember

  • AR(1) is stationary when ∣ϕ∣<1|\phi| < 1∣ϕ∣<1; ϕ=1\phi = 1ϕ=1 is a random walk.

Unit roots, spurious regression and the basis of pairs trading

The unit root

xt=ϕ xt−1+εtx_t = \phi\,x_{t-1} + \varepsilon_txt​=ϕxt−1​+εt​

Stationary when ∣ϕ∣<1|\phi| < 1∣ϕ∣<1; a random walk when ϕ=1\phi = 1ϕ=1. That single boundary separates a series that mean-reverts from one that does not.

The augmented Dickey–Fuller regression

Δyt=α+γ yt−1+∑i=1pδi Δyt−i+εt\Delta y_t = \alpha + \gamma\,y_{t-1} + \sum_{i=1}^{p}\delta_i\,\Delta y_{t-i} + \varepsilon_tΔyt​=α+γyt−1​+i=1∑p​δi​Δyt−i​+εt​

Test γ=0\gamma = 0γ=0 (a unit root) against γ<0\gamma < 0γ<0 (mean reversion). The lagged differences soak up short-run autocorrelation; the test statistic is compared with Dickey–Fuller critical values, which are more negative than the usual ttt ones — around −2.9-2.9−2.9 at 5%5\%5% with a constant.

Remember

  • Prices have unit roots; returns are usually stationary.

Volatility models: ARCH, GARCH and realised measures

GARCH(1,1)

σt2=ω+αεt−12+βσt−12\sigma_t^2 = \omega + \alpha\varepsilon_{t-1}^2 + \beta\sigma_{t-1}^2σt2​=ω+αεt−12​+βσt−12​

Today’s variance is a constant, plus a reaction to yesterday’s squared shock, plus a memory of yesterday’s variance.

GJR-GARCH

σt2=ω+(α+γ 1{εt−1<0})εt−12+βσt−12\sigma_t^2 = \omega + \big(\alpha + \gamma\,\mathbf{1}\{\varepsilon_{t-1} < 0\}\big)\varepsilon_{t-1}^2 + \beta\sigma_{t-1}^2σt2​=ω+(α+γ1{εt−1​<0})εt−12​+βσt−12​

An extra coefficient on negative shocks captures the leverage effect. Persistence becomes α+γ/2+β\alpha + \gamma/2 + \betaα+γ/2+β for symmetric shocks. In equity indices α\alphaα is often close to zero and nearly all the reaction comes through γ\gammaγ.

Remember

  • σt2=ω+αεt−12+βσt−12\sigma_t^2 = \omega + \alpha\varepsilon_{t-1}^2 + \beta\sigma_{t-1}^2σt2​=ω+αεt−12​+βσt−12​.

State space and the Kalman filter

State-space form

xt=Fxt−1+wt⏟state,yt=Hxt+vt⏟observation\underbrace{x_t = Fx_{t-1} + w_t}_{\text{state}}, \qquad \underbrace{y_t = Hx_t + v_t}_{\text{observation}}statext​=Fxt−1​+wt​​​,observationyt​=Hxt​+vt​​​

The state evolves on its own and you observe a noisy function of it. Nearly every time-series model can be written this way, which is what makes the filter so general.

The local linear trend model

yt=μt+εt,μt+1=μt+νt+ηt,νt+1=νt+ζty_t = \mu_t + \varepsilon_t,\quad \mu_{t+1} = \mu_t + \nu_t + \eta_t,\quad \nu_{t+1} = \nu_t + \zeta_tyt​=μt​+εt​,μt+1​=μt​+νt​+ηt​,νt+1​=νt​+ζt​

A level that drifts with a slowly changing slope, observed with noise. Setting the slope noise to zero gives a random walk with fixed drift; setting both to zero gives a straight line. It is the natural state-space model for a trending signal whose trend itself evolves.

The likelihood comes free

ln⁡L=−12∑t(ln⁡(2πFt)+vt2Ft),vt=yt−Hmt−, Ft=HPt−H⊤+R\ln L = -\tfrac{1}{2}\sum_t\left(\ln(2\pi F_t) + \frac{v_t^2}{F_t}\right),\quad v_t = y_t - H m_t^-,\ F_t = HP_t^-H^\top + RlnL=−21​t∑​(ln(2πFt​)+Ft​vt2​​),vt​=yt​−Hmt−​, Ft​=HPt−​H⊤+R

The prediction errors (innovations) and their variances from the filter give the exact Gaussian likelihood, so the unknown variances QQQ and RRR can be estimated by maximising it. That is how the signal-to-noise ratio is chosen from data rather than by hand.

Remember

  • State space separates a hidden state from noisy observations of it.

Stylised facts: what financial returns actually look like

The variance ratio test

VR(q)=Var(rt(q))/qVar(rt)VR(q) = \frac{\mathrm{Var}(r_t^{(q)})/q}{\mathrm{Var}(r_t)}VR(q)=Var(rt​)Var(rt(q)​)/q​

Under a random walk, variance grows linearly with the horizon, so the ratio is one. Above one means trending; below means mean reversion.

Remember

  • Fat tails, no return autocorrelation, volatility clustering, leverage effect.

Forecast evaluation: walk-forward design and backtest overfitting

Diebold–Mariano

DM=dˉVar^(dˉ),dt=L(e1t)−L(e2t)DM = \frac{\bar{d}}{\sqrt{\widehat{\mathrm{Var}}(\bar{d})}}, \qquad d_t = L(e_{1t}) - L(e_{2t})DM=Var(dˉ)​dˉ​,dt​=L(e1t​)−L(e2t​)

Test whether two forecasts differ in accuracy by testing whether their loss differential has mean zero.

The expected maximum Sharpe from luck

E[max⁡i≤NSRi]≈2ln⁡NT\mathbb{E}\left[\max_{i \le N} SR_i\right] \approx \frac{\sqrt{2\ln N}}{\sqrt{T}}E[i≤Nmax​SRi​]≈T​2lnN​​

The best of NNN worthless strategies, over TTT years of data. Anything below this is not evidence of anything.

Minimum backtest length

MinBTL≈2ln⁡NSR2 years\text{MinBTL} \approx \frac{2\ln N}{SR^2}\ \text{years}MinBTL≈SR22lnN​ years

Bailey, Borwein, López de Prado and Zhu: roughly how many years of data are needed before the best of NNN independent skill-less strategies is unlikely to show an annual Sharpe ratio of SRSRSR by luck. More trials need longer histories; higher claimed Sharpe needs shorter ones.

Remember

  • Walk-forward, with every fitted choice inside the loop.

Detailed formula cards

  • Mean-reversion half-life

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.