How the algorithm works
1. Catalogue: each item becomes an attribute vector
Nike Air Max: {brand: Nike, category: sneakers,
colour: white, price: 120, material: mesh}
2. User profile: the weighted average of the attributes of viewed items
User_123: {brand: Nike ×0.6, Adidas ×0.3, category: sneakers ×0.8,
price band: 90–160}
3. Recommendation: the items most similar to the profile
→ cosine similarity(User_123, Item_A) = 0.91 ✓
→ cosine similarity(User_123, Item_B) = 0.34 ✗
Content-based versus collaborative filtering
| Dimension | Content-based | Collaborative |
|---|---|---|
| Works from | Item attributes | Behaviour of similar users |
| Cold start | ✓ Works | ✗ Fails |
| New items | ✓ Works | ✗ No data |
| Serendipity | ✗ Predictable | ✓ Finds the unexpected |
| Precision at high data volume | Medium | High |
The hybrid approach
Most production systems combine algorithms:
New visitor (< 5 events): content-based (100%)
Returning (5–20 events): content-based (60%) + collaborative (40%)
Active (20+ events): collaborative (70%) + content-based (30%)
New item (< 10 sales): content-based (80%) + popularity (20%)
The hybrid removes the weakness each algorithm carries on its own.
Where it is used in e-commerce
- Similar items on a product page — the closest attribute matches to the item being viewed
- New arrivals — items with no interaction history that collaborative filtering cannot place
- Niche categories — low-traffic sections where behavioural data never reaches critical mass
- Explainable recommendations — because you looked at X, a message that only attribute logic can
justify honestly
The practical rule is to treat content-based filtering as the floor of a recommender rather than its
ceiling: it guarantees a relevant answer in every situation, and collaborative models add the
precision on top once the data is there.