Inference: from training to prediction
An ML model goes through two fundamentally different phases. Training — the model learns
patterns from historical data and fits its parameters. Inference — the finished model is applied
to new data to produce predictions.
Only inference runs in production. A shopper opens a category page, the system triggers inference on
the recommender model, and within tens of milliseconds a personalized product list comes back.
Inference requirements in e-commerce
Online inference works under hard time constraints:
| Scenario | Acceptable latency | Reason |
|---|---|---|
| Recommendations on a product page | < 50–100 ms | Blocks page rendering |
| Category personalization | < 100 ms | Loads together with the product list |
| AI shopping assistant | < 2–3 s | The shopper is waiting for an answer |
| Batch scoring of segments | No constraint | Offline processing |
Breaching the latency budget in the first two scenarios means a fallback — showing default,
non-personalized recommendations — or a rendering delay, and both affect conversion.
Online versus batch inference
Online inference is synchronous: a request arrives, the model runs, the answer returns. It
serves personalization, recommendations and real-time search, and it needs highly available
inference infrastructure.
Batch inference is asynchronous: the model periodically processes large volumes of data and
writes the results to a database. When a page is requested, the system reads the precomputed result
— fast, but the data may be stale.
In practice recommender systems usually combine both: batch for long-term preferences (an affinity
profile recomputed every few hours) plus online for short-term signals from the current session.
Tip: cache inference results for the standard scenarios such as bestsellers and trending. That
cuts load on the inference server by 60 to 80% without sacrificing quality for most shoppers.
LLM inference: a special case
For large language models, inference is an order of magnitude more resource-hungry than for
classical ML models. A single token in the answer of a GPT-class model takes gigaflop-scale
computation. Hence:
- Dedicated GPU or TPU servers for inference
- Quantisation — lowering numeric precision for speed with little quality loss
- Streaming output — tokens are sent to the user as they are generated, instead of waiting for the
full answer - KV-cache reuse to speed up repeated context