How gradient boosting works

A classical machine learning algorithm builds one model at once. Boosting works differently: it
assembles an ensemble of weak models that correct each other’s errors in sequence.

The process:
1. The first tree is built as a rough approximation of the target variable
2. Residuals are computed — how far the tree missed on each object
3. The next tree is trained to predict those residuals
4. The final prediction is a weighted sum of all the trees

The gradient in the name comes from the fact that the direction of each next tree is set by the
gradient of the loss function, which generalises the method to any differentiable loss, not only MSE.

Final prediction = T1(x) + eta * T2(x) + eta * T3(x) + ...
where eta is the learning rate (0.01 to 0.3)

Important: the lower the learning rate, the more trees you need — and the more resistant the
model is to overfitting. The standard practice is a learning rate of 0.05 to 0.1 plus early
stopping on a validation set.

Applications in e-commerce

Predictive analytics

Boosting is the de facto standard for binary classification and regression on tabular data:
– Churn prediction
– Scoring the probability of a purchase within a session
– Predicting the probability of a return
– Dynamic credit risk assessment (BNPL, instalment plans)

Ranking inside recommendations

In two-stage recommender pipelines, boosting occupies the second stage — reranking. Once an
embedding model (two-tower, Item2Vec) has selected 100 to 500 candidates, LightGBM or XGBoost
reorders them using:

Feature Description
Session context Items viewed, category of the current page
Customer attributes Segment, RFM, purchase history
Product attributes Margin, availability, rating, newness
Merchandising rules Boost for priority positions

Search ranking

In e-commerce search, boosting ranks results by learning from clicks and purchases — learning to
rank, with LambdaMART being the boosting variant built for ranking objectives.

Regularisation and overfitting

The main parameters that guard against overfitting:
– max_depth — tree depth (3 to 6 is enough for boosting; deep trees are not needed)
– min_child_samples — the minimum number of objects in a leaf
– subsample / colsample_bytree — random subsampling of rows and features
– reg_alpha, reg_lambda — L1 and L2 regularisation of leaf weights

Tuned properly and fed well-prepared features, gradient boosting competes with neural networks on
most tabular e-commerce tasks.