What an embedding is
A computer does not know that sneakers and running shoes are related concepts. For it to find
semantically similar objects, they have to be represented as numbers — and in a way that puts
similar meanings close together in the number space.
An embedding is a vector, a list of numbers, encoding an object:
Nike Air Max sneakers → [0.23, -0.11, 0.87, ..., 0.04] # 128 numbers
Adidas running shoes → [0.21, -0.09, 0.91, ..., 0.07] # close
iPhone smartphone → [-0.85, 0.62, -0.31, ..., 0.55] # far
The distance between vectors — cosine or Euclidean — reflects semantic similarity. Finding similar
products becomes finding the vectors nearest to the current item’s vector.
How embeddings are trained in e-commerce
From behaviour (Item2Vec). Session sequences are the training data: a shopper viewed [A, B, C,
D]. The model learns to predict the context, so items viewed together receive close vectors. This
behavioural similarity captures patterns invisible in the descriptions.
From text. An LLM or a specialised model encodes the product description into a vector. This
suits semantic search: a query such as lightweight shoes for summer runs finds relevant products
even when that exact wording never appears in a description.
A two-tower model. Separate embeddings for users and for items. Closeness between a user vector
and an item vector is personalized relevance. This is the backbone of modern recommenders.
Applications in e-commerce
| Scenario | What is encoded | Task |
|---|---|---|
| Similar items | Products | Nearest-neighbour lookup by vector |
| Semantic search | Query plus products | Matching a query to the catalogue |
| Personal recommendations | Users plus products | Two-tower matching |
| Interest profile | Behaviour history | The mean vector of views as a taste profile |
Tip: the mean vector of a shopper’s viewed items is a simple and effective taste profile. The
nearest items they have not seen yet make strong recommendations at cold start.
Common mistakes
- Training on clicks alone. Clicks are a noisy signal. Fine-tune on purchases or add-to-cart
events, weighting by signal strength. - One embedding set for search and recommendations. These are different tasks: search optimises
relevance to a query, recommendations optimise engagement and conversion. - No retraining. The catalogue changes and behaviour moves. Embeddings trained six months ago do
not reflect seasonal trends or new arrivals.