The role of the catalogue in personalization

The product catalogue is the entry point of any recommender system. Algorithms do not work from
images and prose but from structured attributes: category_id, brand, price, in_stock,
color. The more complete that data, the more precisely an algorithm can find similar products and
match them to a shopper’s interests.

What goes into a catalogue

The minimum attribute set a personalization platform needs:

Attribute Why it is needed
item_id / SKU The unique product identifier
category / subcategory Classification for content-based algorithms
brand Behavioural brand loyalty
price The shopper’s price band
in_stock Filtering unavailable items out
image_url Rendering inside the widget
product_url The link to click through to

Extended attributes raise the precision of content-based and search algorithms: colour,
material, gender, age group, rating, tags and technical specifications.

Keeping the catalogue fresh

A stale catalogue damages the experience directly. If a recommended item is out of stock, the
shopper lands on a 404 or on a page with no way to buy.

Recommended refresh rates:

  • in_stock and price — every 30–60 minutes, or on an event when the value changes
  • Item attributes — once a day
  • New items — immediately, through the Event API

Catalogue coverage

High coverage is an indicator of a healthy recommender system. If the algorithm only ever
recommends popular products, the store loses the revenue sitting in the long tail. The levers for
raising coverage are recommendation diversity, boost rules for under-exposed categories and
explorative New arrivals blocks.