How MAP is computed
MAP is built on two levels of aggregation.
Step 1 — average precision for a single user:
Picture a list of 5 recommendations in which positions 1, 3 and 5 are relevant:
Position: 1 2 3 4 5
Hit: yes no yes no yes
P@1 = 1/1 = 1.0 (the 1st relevant item was found at position 1)
P@3 = 2/3 = 0.67 (the 2nd relevant item was found at position 3)
P@5 = 3/5 = 0.60 (the 3rd relevant item was found at position 5)
AP = (1.0 + 0.67 + 0.60) / 3 = 0.76
Step 2 — MAP is the mean of the AP values across all users or queries.
Why order matters
The difference between two algorithms is often not how many relevant items they find, but where
they place them. A shopper typically scans the first three to five recommendations. An algorithm
that puts relevant products at the top scores higher on MAP — and has a real effect on clicks and
conversion.
MAP versus NDCG: which to choose
| Criterion | MAP | NDCG |
|---|---|---|
| Relevance type | Binary (yes/no) | Graded (1, 2, 3…) |
| Sensitivity to order | High | High |
| Uneven number of relevant items | Handled correctly | Handled correctly |
| Typical use | Search, basic recommendations | Personalized recommendations |
Tip: for offline evaluation of a recommendation engine, compute several metrics at once — MAP,
NDCG@10 and coverage. A single metric can mask problems along the other dimensions.
Practical limitations
MAP assumes that the full list of relevant products is known for every user. In practice that is
hard: interaction history is incomplete, because the shopper simply never saw part of the catalogue.
Offline metrics therefore have to be complemented by online experiments — A/B tests — where the
winner is decided by real conversion and revenue.