How the algorithm works

1. Catalogue: each item becomes an attribute vector
   Nike Air Max: {brand: Nike, category: sneakers,
                  colour: white, price: 120, material: mesh}

2. User profile: the weighted average of the attributes of viewed items
   User_123: {brand: Nike ×0.6, Adidas ×0.3, category: sneakers ×0.8,
              price band: 90–160}

3. Recommendation: the items most similar to the profile
   → cosine similarity(User_123, Item_A) = 0.91 ✓
   → cosine similarity(User_123, Item_B) = 0.34 ✗

Content-based versus collaborative filtering

Dimension Content-based Collaborative
Works from Item attributes Behaviour of similar users
Cold start ✓ Works ✗ Fails
New items ✓ Works ✗ No data
Serendipity ✗ Predictable ✓ Finds the unexpected
Precision at high data volume Medium High

The hybrid approach

Most production systems combine algorithms:

New visitor (< 5 events):      content-based (100%)
Returning (5–20 events):       content-based (60%) + collaborative (40%)
Active (20+ events):           collaborative (70%) + content-based (30%)
New item (< 10 sales):         content-based (80%) + popularity (20%)

The hybrid removes the weakness each algorithm carries on its own.

Where it is used in e-commerce

  • Similar items on a product page — the closest attribute matches to the item being viewed
  • New arrivals — items with no interaction history that collaborative filtering cannot place
  • Niche categories — low-traffic sections where behavioural data never reaches critical mass
  • Explainable recommendations — because you looked at X, a message that only attribute logic can
    justify honestly

The practical rule is to treat content-based filtering as the floor of a recommender rather than its
ceiling: it guarantees a relevant answer in every situation, and collaborative models add the
precision on top once the data is there.