Why overfitting is a practical problem for recommendations

Overfitting is not a theoretical concept but a real cause of recommender degradation in production.
A model overfitted to historical data reproduces the past instead of predicting the future.

In e-commerce it looks like this:
– Recommendations stuck on the bestsellers of six months ago
– The same already-purchased product recommended again and again
– New catalogue items never appearing in the output because the model never saw them in training

The bias-variance tradeoff

Simple model:      high bias      → underfitting
Complex model:     high variance  → overfitting
Optimal model:     a balance between bias and variance

The job of training is to find the sweet spot where the model has captured the real patterns but
has not learned the noise.

How overfitting is controlled

Regularisation — adding a penalty on model complexity (L1, L2, dropout in neural networks). It
prevents individual weights from growing too large.

A correct data split — train, validation and test with no leakage of future data into training.
In recommendations the temporal split matters: train on the past, test on the following period.

Early stopping — halting neural network training once the validation metric stops improving.

Reducing complexity — cutting the number of layers, lowering embedding dimensionality,
simplifying the architecture.

Tip: retrain models on fresh data regularly. Shopper behaviour changes, and a model that has
not been refreshed for several months will degrade inevitably — even without classic overfitting.