The problem: how to measure whether recommendations are good
A recommender returns a list of K products. How do you tell whether that list is any good? Two
metrics answer the question: precision and recall.
Precision answers: “How many of the recommended products turned out to be relevant?”
Recall answers: “How many of all the products relevant to this shopper did the system actually
find and show?”
Example:
Products genuinely relevant to the user: 20
Products the system recommended: 10
Of which relevant: 7
Precision = 7 / 10 = 0.70 (70% of the recommendations hit the target)
Recall = 7 / 20 = 0.35 (35% of everything relevant was covered)
The precision versus recall trade-off
There is a fundamental trade-off between the two. To raise recall — to cover more of the relevant
products — the system has to recommend more items or filter less strictly, and that lets more
irrelevant ones through, so precision falls.
| Scenario | What changes | Effect |
|---|---|---|
| Increase K (more recommendations) | ↑ Recall | ↓ Precision |
| Tighten the score threshold | ↑ Precision | ↓ Recall |
| Show bestsellers only | ↑ Precision | ↓ Recall (the long tail disappears) |
Precision@K and Recall@K
In recommender systems the metrics are always computed at a fixed K — the number of slots in the
widget:
- Precision@5 — for homepage widgets with five slots
- Precision@10 — for a horizontal widget on a product page
- Recall@20 — for search or an email selection
That keeps the metrics comparable and tied to the real conditions of the interface.
F1 score: when you need the balance
The F1 score is the harmonic mean of precision and recall. It is useful when neither has explicit
priority and several models have to be compared with one number.
F1 = 2 × (Precision × Recall) / (Precision + Recall)
Model A: Precision 0.80, Recall 0.80 → F1 = 0.80
Model B: Precision 0.95, Recall 0.20 → F1 = 0.33
Model B looks precise, yet it covers only 20% of what is relevant — F1 states that weakness
honestly.
Tip: precision and recall are offline metrics computed on historical data. To judge a
recommendation algorithm for real they have to be confirmed by online metrics in an A/B test:
widget CTR and attributed revenue.