How latency is measured

Latency is described by a distribution, not a single number. The standard set of metrics:

Metric What it means
p50 (median) Half of all requests finish faster than this
p95 95% of requests finish faster; 5% are slower
p99 99% of requests finish faster; 1% are slower (the tail)
p999 Used for systems with very strict requirements
Latency distribution (ms):
p50  = 12 ms   <- a typical fast request
p95  = 38 ms   <- 5% slightly slower
p99  = 87 ms   <- the tails: cache misses, GC pauses
p999 = 420 ms  <- rare outliers

A personalization platform’s SLA is usually expressed at p99 — the guaranteed behaviour of the system excluding the rarest anomalies.

Where latency comes from in a personalization API

The total is the sum of several components:

  1. Network round trip (RTT) — the time to the data centre and back. Reduced with a CDN and by placing capacity closer to users.
  2. Server queueing — at peak load requests wait for a worker to free up.
  3. Computing the recommendations — a nearest-neighbour search in a vector store or a ranking model pass.
  4. Data access — reading the user profile, history and session context from storage.
  5. Response serialisation — building the JSON and compressing it.

Tip: cache precomputed recommendations for popular scenarios — top products, frequently browsed categories. Recompute personalized recommendations asynchronously rather than on the critical path of page load; that cuts p99 by an order of magnitude.

Latency and Core Web Vitals

During page load, personalization API latency feeds directly into LCP (Largest Contentful Paint) whenever the recommendation widget lands in the critical render. Recommendation blocks in the footer or below the fold do not affect LCP, which is why lazy loading widgets is standard practice for Core Web Vitals optimisation when working with personalization platforms.