Decomposing a model’s error

The expected prediction error breaks into three components:

Error = Bias^2 + Variance + Irreducible noise
  • Bias — the systematic deviation of predictions from the true values. The cause is assumptions
    about the data that are too simple.
  • Variance — the sensitivity of predictions to the specific training sample. The cause is a
    model complex enough to have memorised the noise.
  • Irreducible noise — the random component of the data that no model can predict.

Visualising the tradeoff

Error
  |          Total error
  |        \             /
  |         \    min   /
  |  Bias^2  \       /  Variance
  |            \   /
  |             \ /
  +-------------------- Model complexity
   Simple             Complex

The optimum sits where total error is lowest — neither at maximum nor at minimum complexity.

Applying it in recommendations

Situation Problem Fix
A linear model misses the patterns High bias Move to matrix factorisation or a two-tower model
The model is excellent on history, poor on new data High variance Strengthen regularisation, add data
Rare items are predicted badly High variance on few observations A content-based fallback for the cold start

Ensembling as a balance

Ensemble methods — Random Forest, Gradient Boosting — work directly on this tradeoff:

  • Bagging (Random Forest): trains many trees on subsamples and averages them, which lowers
    variance without raising bias
  • Boosting (XGBoost, LightGBM): corrects errors in sequence, which lowers bias while variance is
    held in check by regularisation

Tip: in recommender systems, mixing strategies — combining popularity, collaborative filtering
and content-based models — is ensembling by another name. Each model has its own bias-variance
profile, and blending cancels out the weaknesses of each.