Standard errors for autocorrelations beyond white noise
Var(ρ^h)≈T1(1+2k=1∑qρk2),h>q
If the true process is MA(q), sample autocorrelations beyond lag q are noisier than 1/T suggests, because earlier dependence propagates into their estimates. Using the white-noise band for every lag finds spurious structure in persistent series.
Remember
Weak stationarity: constant mean and variance, covariance depending only on the lag.
AR terms carry the series’ own memory; MA terms carry the memory of past shocks.
How much of the variance a forecast leaves
Var(xt)Var(et+h)=1−ϕ2h(AR(1))
The fraction of the unconditional variance that remains unexplained h steps ahead. At h=1 it is 1−ϕ2; as h grows it approaches one, and the forecast becomes the mean.
Persistence is underestimated in short samples
E[ϕ^]≈ϕ−T1+3ϕ
The least-squares estimate of an AR(1) coefficient, with an estimated mean, is biased toward zero by roughly this amount (Kendall; Marriott and Pope). The bias is largest precisely for persistent series.
Remember
AR(1) is stationary when ∣ϕ∣<1; ϕ=1 is a random walk.
Stationary when ∣ϕ∣<1; a random walk when ϕ=1. That single boundary separates a series that mean-reverts from one that does not.
The augmented Dickey–Fuller regression
Δyt=α+γyt−1+i=1∑pδiΔyt−i+εt
Test γ=0 (a unit root) against γ<0 (mean reversion). The lagged differences soak up short-run autocorrelation; the test statistic is compared with Dickey–Fuller critical values, which are more negative than the usual t ones — around −2.9 at 5% with a constant.
Remember
Prices have unit roots; returns are usually stationary.
Today’s variance is a constant, plus a reaction to yesterday’s squared shock, plus a memory of yesterday’s variance.
GJR-GARCH
σt2=ω+(α+γ1{εt−1<0})εt−12+βσt−12
An extra coefficient on negative shocks captures the leverage effect. Persistence becomes α+γ/2+β for symmetric shocks. In equity indices α is often close to zero and nearly all the reaction comes through γ.
The state evolves on its own and you observe a noisy function of it. Nearly every time-series model can be written this way, which is what makes the filter so general.
The local linear trend model
yt=μt+εt,μt+1=μt+νt+ηt,νt+1=νt+ζt
A level that drifts with a slowly changing slope, observed with noise. Setting the slope noise to zero gives a random walk with fixed drift; setting both to zero gives a straight line. It is the natural state-space model for a trending signal whose trend itself evolves.
The prediction errors (innovations) and their variances from the filter give the exact Gaussian likelihood, so the unknown variances Q and R can be estimated by maximising it. That is how the signal-to-noise ratio is chosen from data rather than by hand.
Remember
State space separates a hidden state from noisy observations of it.
Test whether two forecasts differ in accuracy by testing whether their loss differential has mean zero.
The expected maximum Sharpe from luck
E[i≤NmaxSRi]≈T2lnN
The best of N worthless strategies, over T years of data. Anything below this is not evidence of anything.
Minimum backtest length
MinBTL≈SR22lnNyears
Bailey, Borwein, López de Prado and Zhu: roughly how many years of data are needed before the best of N independent skill-less strategies is unlikely to show an annual Sharpe ratio of SR by luck. More trials need longer histories; higher claimed Sharpe needs shorter ones.
Remember
Walk-forward, with every fitted choice inside the loop.