What evals are
Language model output is non-deterministic: the same prompt edit can improve some conversations and quietly break others. Evals are offline quality checks before release: a reference set of conversations (a golden set) is run through the new prompt, model or retrieval version, and metrics are scored. They are the equivalent of automated tests in software, except that you do not check for a word-for-word match — you check against criteria: the facts are correct, the products fit, the rules are respected.
What an eval is made of
Reference set. Real shopper questions plus verifiable expectations for the answer: which products are acceptable, which facts are mandatory, what must not be said. It must include hard cases: the product is out of stock, the question is outside the assortment, a prompt injection attempt.
Metrics:
| Metric | What it shows | Who scores it |
|---|---|---|
| Factual accuracy | Prices, specifications and availability match the catalogue | Code: checked against the feed |
| Faithfulness | The answer rests on the supplied context — see grounding | LLM judge |
| Relevance | The suggested products fit the request | LLM judge, labelling, precision and recall |
| Rule compliance | No discount promises or off-limits topics — a check on guardrails | Code and a classifier |
| Refusal rate | How often the assistant says “I don’t know” or hands over to a human | A counter: a rise signals over-tight rules |
Scorers. Deterministic code for anything that can be checked against data. LLM-as-a-judge — a separate model with a rubric — for semantic criteria. Human labelling to calibrate the judge and settle disputed cases.
Evals and the A/B test
Evals answer “did we break anything?”; an A/B test answers “is it better for the business?”. The workflow:
- A change: a new prompt, model or retrieval setting.
- Evals as a gate: a version with lower factual accuracy or new rule violations goes no further.
- An A/B test on live traffic, measured on CR, AOV and revenue per visit.
- The winner is rolled out and the errors found are added to the reference set.
Important: a high judge score does not guarantee higher conversion. Evals do not replace A/B tests; they save traffic by stopping weak versions before shoppers ever see them.
Getting started
- Pull the first few dozen real conversations from the logs and write verifiable expectations for each, not a reference answer text.
- Automate the run on every change to the prompt, model or retrieval, and on every major catalogue update.
- Set thresholds that block a release: any rule violation, any drop in factual accuracy.
- Once per cycle, check the LLM judge against human labels on a sample of conversations.
- Turn every error found in production into a new test case.