What evals are

Language model output is non-deterministic: the same prompt edit can improve some conversations and quietly break others. Evals are offline quality checks before release: a reference set of conversations (a golden set) is run through the new prompt, model or retrieval version, and metrics are scored. They are the equivalent of automated tests in software, except that you do not check for a word-for-word match — you check against criteria: the facts are correct, the products fit, the rules are respected.

What an eval is made of

Reference set. Real shopper questions plus verifiable expectations for the answer: which products are acceptable, which facts are mandatory, what must not be said. It must include hard cases: the product is out of stock, the question is outside the assortment, a prompt injection attempt.

Metrics:

Metric What it shows Who scores it
Factual accuracy Prices, specifications and availability match the catalogue Code: checked against the feed
Faithfulness The answer rests on the supplied context — see grounding LLM judge
Relevance The suggested products fit the request LLM judge, labelling, precision and recall
Rule compliance No discount promises or off-limits topics — a check on guardrails Code and a classifier
Refusal rate How often the assistant says “I don’t know” or hands over to a human A counter: a rise signals over-tight rules

Scorers. Deterministic code for anything that can be checked against data. LLM-as-a-judge — a separate model with a rubric — for semantic criteria. Human labelling to calibrate the judge and settle disputed cases.

Evals and the A/B test

Evals answer “did we break anything?”; an A/B test answers “is it better for the business?”. The workflow:

  1. A change: a new prompt, model or retrieval setting.
  2. Evals as a gate: a version with lower factual accuracy or new rule violations goes no further.
  3. An A/B test on live traffic, measured on CR, AOV and revenue per visit.
  4. The winner is rolled out and the errors found are added to the reference set.

Important: a high judge score does not guarantee higher conversion. Evals do not replace A/B tests; they save traffic by stopping weak versions before shoppers ever see them.

Getting started

  • Pull the first few dozen real conversations from the logs and write verifiable expectations for each, not a reference answer text.
  • Automate the run on every change to the prompt, model or retrieval, and on every major catalogue update.
  • Set thresholds that block a release: any rule violation, any drop in factual accuracy.
  • Once per cycle, check the LLM judge against human labels on a sample of conversations.
  • Turn every error found in production into a new test case.