What machine learning is

A classical program executes rules written by a person. A machine learning model receives historical
examples and adjusts its parameters to minimise prediction error on them — then applies the
relationship it found to new data.

In e-commerce the difference shows in a simple example. The rule show accessories from the same
brand encodes one hypothesis from one category manager. A model trained on millions of sessions
discovers that in one category the decisive factor is the price band, in another the combination of
brand and size, in a third the time since the previous purchase. None of those were specified.

That does not make rules unnecessary. Industrial systems are almost always hybrid: the model handles
ranking under uncertainty, the rules handle business constraints the model cannot know — supplier
contracts, a collection clearance, legal restrictions.

Types of learning

Type Input E-commerce tasks
Supervised Examples with known correct answers Churn prediction, response scoring, ranking
Unsupervised Unlabelled data Customer segmentation, similar items, anomaly detection
Reinforcement Feedback from the environment Dynamic traffic allocation, output optimisation

Supervised learning is the main working tool: almost any business task reduces to predicting a
number or a class from a feature set. Unsupervised learning is used where labels do not exist at all
— clustering a customer base into behaviourally similar groups, for instance. Reinforcement learning
appears in retail mostly in its light form, the multi-armed bandit, where the algorithm gradually
shifts traffic toward the winning variation.

A separate branch is deep learning on neural networks. It produces vector representations of objects
and works well with behavioural sequences and text, but demands more data and infrastructure. For
tabular problems, gradient boosting remains the practical choice on quality per unit of cost.

Features matter more than the algorithm

Model quality is determined first by the features available to it. Feature engineering turns raw
events into numbers describing an object.

Raw event:
  {user: u_18422, event: product_view, sku: 77301, ts: 2026-03-14T19:22:10}

Features built from the event stream:
  views in 7 days                      = 34
  unique categories in 30 days         = 6
  mean order value of last 3 orders    = 96.40
  days since last purchase             = 41
  share of views in the 60–120 band    = 0.62
  affinity to brand A                  = 0.31

The mistake that invalidates all of it is target leakage: a feature carrying information unavailable
at prediction time. A feature such as total orders this month, used to predict a purchase in that
same month, looks excellent in validation and is worthless in production. The rule is simple: every
feature is computed strictly from data preceding the moment being predicted.

The model lifecycle

1. Data collection  → site and app events, catalogue, orders
2. Features         → aggregates per user, per item, per user × item pair
3. Training         → fitting parameters on the training set
4. Validation       → checking on a period held out in time
5. Offline scoring  → comparison with the current solution
6. A/B test         → verification on live traffic against a control
7. Inference        → serving in production inside a latency budget
8. Monitoring       → quality metrics, feature distributions, drift
9. Refresh          → regular retraining on current data

Two points where results are most often lost.

Validation. A random split on data with temporal structure inflates the estimate: the model
peeks into the future. The correct scheme trains on the period before date X and validates after it.
Overfitting is caught exactly here, and regularisation and model simplification are the standard
responses.

Inference. A model needing 300 ms to answer cannot rank a listing. Real-time inference fits into
tens of milliseconds, so heavy computation is precomputed: item embeddings and user aggregates are
prepared in advance, leaving a fast dot product and a sort online. Some tasks need no online path at
all — churn prediction runs as a nightly batch.

⚠

A model degrades without a single failure. The assortment refreshes, traffic structure changes, the
site’s event schema is edited, and the feature distribution drifts away from the one used in
training. Monitoring input distributions is as mandatory as monitoring quality metrics: it catches
the problem before the business sees it.

ML tasks in e-commerce

Task Learning type Output How it is verified
Product recommendations Supervised plus unsupervised A ranked SKU list A/B test on revenue per visitor
Search and listing ranking Supervised (learning to rank) The order of results A/B test on category conversion
Customer segmentation Unsupervised A partition of the base Interpretability plus campaign response
Churn prediction Supervised (classification) A churn probability per customer Held-out quality plus retention campaign effect
Similar items Unsupervised (embeddings) Vector proximity of SKUs Clicks and add-to-cart from the block
Demand anomalies Unsupervised A deviation flag Manual verification by a category manager
Traffic allocation in tests Reinforcement Traffic shares per variation Total experiment revenue

The boundary of applicability is worth stating. Machine learning predicts what is encoded in
historical data. If an event is not logged, the model does not know about it; if the process changed
discontinuously — a new market, a new business model — history says little about the future.

A checklist before starting an ML project

  1. State the prediction in terms of a business decision. Predict churn probability is useless
    without an answer to what action follows a high score.
  2. Confirm the target event is logged and the example count is sufficient — count examples of
    the target class, not total data volume.
  3. Fix a baseline. A simple rule or popularity is the mandatory reference point.
  4. Validate by time, not by random split.
  5. Check features for target leakage.
  6. Set the latency budget before choosing an architecture, not after.
  7. Provide a fallback for when the model is unavailable — non-personal output beats an empty block.
  8. Plan regular retraining and drift monitoring from the first day in production.
  9. Verify the effect with an A/B test. An offline metric improvement is not a result until live
    traffic confirms it.