The model
y=Xβ+ε,E[ε∣X]=0 Linear in the parameters — not necessarily in the variables, since x2 and logx are perfectly acceptable columns.
The one-regressor case
β^=Var(x)Cov(x,y)=ρσxσy Worth memorising in the second form: a beta is a correlation scaled by the ratio of standard deviations.
The variance of the estimator
Var(β^∣X)=σ2(X⊤X)−1 Under homoskedastic, uncorrelated errors. The diagonal gives each coefficient’s squared standard error; the off-diagonals say how their estimation errors move together. Correlated regressors make X⊤X nearly singular and blow up the diagonal of its inverse.
Remember
- β^=(X⊤X)−1X⊤y, and the residual is orthogonal to every regressor.
R squared and its adjustment
R2=1−TSSRSS,Rˉ2=1−TSS/(n−1)RSS/(n−k−1) The share of variance explained, and the version that charges for parameters.
The F test for nested models
F=RSSu/(n−k−1)(RSSr−RSSu)/q Tests whether the q extra regressors in the unrestricted model add anything jointly.
What one more regressor adds
Rnew2=Rold2+ryxk⋅rest2(1−Rold2) A new regressor explains a share of what was left unexplained, equal to its squared partial correlation with y given the existing regressors. It can only raise R2, and it raises it by little when the model already explains most of the variance.
Remember
- R2 never falls when you add a regressor; adjusted R2 can.
Logistic regression
ln1−pp=β0+β⊤x⟺p=1+e−(β0+β⊤x)1 Model the log-odds linearly, which maps the whole real line onto (0,1) so the fitted probability is always valid.
The quantile loss
ρτ(u)=u(τ−1{u<0}) Quantile regression minimises this asymmetric loss. For the median (τ=0.5) it is half the absolute error; for a low quantile it charges little for over-predicting and a lot for under-predicting, which pushes the fit down to the chosen quantile.
Remember
- Logistic models the log-odds; eβ is the odds multiplier.