The framework: bias, variance, capacity and dimensionality
The error decomposition
Three terms, of which only the first two are under your control.
How much worse a model does out of sample
For least squares with predictors, observations, noise variance and Gaussian regressors, when the true coefficients are all zero. Training error flatters the model by the fraction ; test error is inflated by estimation noise in every coefficient.
Remember
- Error is bias squared plus variance plus irreducible noise.