Decomposing a model’s error
The expected prediction error breaks into three components:
Error = Bias^2 + Variance + Irreducible noise
- Bias — the systematic deviation of predictions from the true values. The cause is assumptions
about the data that are too simple. - Variance — the sensitivity of predictions to the specific training sample. The cause is a
model complex enough to have memorised the noise. - Irreducible noise — the random component of the data that no model can predict.
Visualising the tradeoff
Error
| Total error
| \ /
| \ min /
| Bias^2 \ / Variance
| \ /
| \ /
+-------------------- Model complexity
Simple Complex
The optimum sits where total error is lowest — neither at maximum nor at minimum complexity.
Applying it in recommendations
| Situation | Problem | Fix |
|---|---|---|
| A linear model misses the patterns | High bias | Move to matrix factorisation or a two-tower model |
| The model is excellent on history, poor on new data | High variance | Strengthen regularisation, add data |
| Rare items are predicted badly | High variance on few observations | A content-based fallback for the cold start |
Ensembling as a balance
Ensemble methods — Random Forest, Gradient Boosting — work directly on this tradeoff:
- Bagging (Random Forest): trains many trees on subsamples and averages them, which lowers
variance without raising bias - Boosting (XGBoost, LightGBM): corrects errors in sequence, which lowers bias while variance is
held in check by regularisation
Tip: in recommender systems, mixing strategies — combining popularity, collaborative filtering
and content-based models — is ensembling by another name. Each model has its own bias-variance
profile, and blending cancels out the weaknesses of each.