How it works

In supervised learning the model receives a dataset of pairs: an input feature vector X and a
target value y. The task is to learn a function f(X) → y that minimises prediction error on new
data.

In e-commerce that looks like this:
– X — purchase history, views, demographics, time since the last purchase
– y — the purchase itself (1/0), churn probability (0–1), the expected order value

Input (features): [7 purchases in 90 days, last one 14 days ago, 3 categories, AOV $25]
Target value:     churn = 0  (did not leave within the next 30 days)

The two main task types

Classification — predicting a categorical answer (yes/no, class A/B/C). Examples: will buy or
will not buy, will churn or will stay, search intent transactional or informational.

Regression — predicting a numeric value. Examples: expected LTV, the forecast order value of the
next purchase, the probability of a return.

Use in personalization

Recommendation algorithms built on supervised learning are trained to predict the probability of an
interaction — a click, a purchase — for a user-item pair:

Task Features (X) Label (y)
Recommendations User profile plus product attributes Click or purchase
Churn prediction RFM features plus behaviour Churn within the next 30 days
PLP ranking User plus position plus product CTR or CR

A critical dependency on data

Label quality determines model quality. The typical problems in e-commerce:

  • Sample bias: the model is trained only on purchased products and never sees the items a
    shopper viewed and abandoned because the page was poor
  • Data leakage: features accidentally contain data from after the target event
  • Class imbalance: a purchase happens in 2% to 3% of cases, so the model takes the lazy route
    and stops predicting the rare class