The problem regularization solves
Machine learning optimises a loss function on the training sample. With no constraints the model
pushes the error towards zero and starts memorising the training data instead of finding patterns
that generalise. On new data such a model performs badly: that is overfitting.
Regularization adds a penalty for model complexity directly into the loss function:
Loss_total = Loss_data + λ × Penalty(weights)
The parameter lambda controls the balance between accuracy on the training set and simplicity of
the model.
The main methods
L1 regularization (lasso):
Penalty = λ × Σ|wᵢ|
Produces sparse models — unimportant weights go to zero. Used for automatic feature selection.
L2 regularization (ridge):
Penalty = λ × Σwᵢ²
Shrinks every weight without zeroing any. More stable when features are multicollinear.
Elastic net: a combination of L1 and L2. Often used when both sparsity and stability are
needed.
Dropout (neural networks):
During training: switch a neuron off with probability p (usually 0.2–0.5)
During inference: all neurons active, weights scaled by (1-p)
Regularization in recommender systems
| Model | Regularization method | Why |
|---|---|---|
| Matrix factorization (ALS, SVD) | L2 on user and item embeddings | Stop it memorising active users |
| Neural CF models | Dropout + L2 | Generalisation to cold-start users |
| Gradient boosting (XGBoost, LightGBM) | L1/L2 + min_child_weight | Control of tree depth |
| Logistic regression in ranking | L1 or elastic net | Selection of the relevant features |
Tip: in e-commerce recommenders it matters especially to regularise the embeddings of active
shoppers. Users with thousands of purchases are heavy examples that a model tends to memorise.
Without regularization it generalises worse to ordinary shoppers with 10 to 50 transactions.
Tuning the lambda hyperparameter
The only correct way is cross-validation:
- Split the data into train / validation / test.
- Train the model with several lambda values (a logarithmic grid: 0.001, 0.01, 0.1, 1.0, 10.0).
- Pick the lambda with the best metric on validation.
- Evaluate on test at the end — once.
Never tune lambda on the test set: that is data leakage.