Inference: from training to prediction

An ML model goes through two fundamentally different phases. Training — the model learns
patterns from historical data and fits its parameters. Inference — the finished model is applied
to new data to produce predictions.

Only inference runs in production. A shopper opens a category page, the system triggers inference on
the recommender model, and within tens of milliseconds a personalized product list comes back.

Inference requirements in e-commerce

Online inference works under hard time constraints:

Scenario Acceptable latency Reason
Recommendations on a product page < 50–100 ms Blocks page rendering
Category personalization < 100 ms Loads together with the product list
AI shopping assistant < 2–3 s The shopper is waiting for an answer
Batch scoring of segments No constraint Offline processing

Breaching the latency budget in the first two scenarios means a fallback — showing default,
non-personalized recommendations — or a rendering delay, and both affect conversion.

Online versus batch inference

Online inference is synchronous: a request arrives, the model runs, the answer returns. It
serves personalization, recommendations and real-time search, and it needs highly available
inference infrastructure.

Batch inference is asynchronous: the model periodically processes large volumes of data and
writes the results to a database. When a page is requested, the system reads the precomputed result
— fast, but the data may be stale.

In practice recommender systems usually combine both: batch for long-term preferences (an affinity
profile recomputed every few hours) plus online for short-term signals from the current session.

Tip: cache inference results for the standard scenarios such as bestsellers and trending. That
cuts load on the inference server by 60 to 80% without sacrificing quality for most shoppers.

LLM inference: a special case

For large language models, inference is an order of magnitude more resource-hungry than for
classical ML models. A single token in the answer of a GPT-class model takes gigaflop-scale
computation. Hence:

  • Dedicated GPU or TPU servers for inference
  • Quantisation — lowering numeric precision for speed with little quality loss
  • Streaming output — tokens are sent to the user as they are generated, instead of waiting for the
    full answer
  • KV-cache reuse to speed up repeated context