What an embedding is

A computer does not know that sneakers and running shoes are related concepts. For it to find
semantically similar objects, they have to be represented as numbers — and in a way that puts
similar meanings close together in the number space.

An embedding is a vector, a list of numbers, encoding an object:

Nike Air Max sneakers    → [0.23, -0.11, 0.87, ..., 0.04]  # 128 numbers
Adidas running shoes     → [0.21, -0.09, 0.91, ..., 0.07]  # close
iPhone smartphone        → [-0.85, 0.62, -0.31, ..., 0.55]  # far

The distance between vectors — cosine or Euclidean — reflects semantic similarity. Finding similar
products becomes finding the vectors nearest to the current item’s vector.

How embeddings are trained in e-commerce

From behaviour (Item2Vec). Session sequences are the training data: a shopper viewed [A, B, C,
D]. The model learns to predict the context, so items viewed together receive close vectors. This
behavioural similarity captures patterns invisible in the descriptions.

From text. An LLM or a specialised model encodes the product description into a vector. This
suits semantic search: a query such as lightweight shoes for summer runs finds relevant products
even when that exact wording never appears in a description.

A two-tower model. Separate embeddings for users and for items. Closeness between a user vector
and an item vector is personalized relevance. This is the backbone of modern recommenders.

Applications in e-commerce

Scenario What is encoded Task
Similar items Products Nearest-neighbour lookup by vector
Semantic search Query plus products Matching a query to the catalogue
Personal recommendations Users plus products Two-tower matching
Interest profile Behaviour history The mean vector of views as a taste profile

Tip: the mean vector of a shopper’s viewed items is a simple and effective taste profile. The
nearest items they have not seen yet make strong recommendations at cold start.

Common mistakes

  • Training on clicks alone. Clicks are a noisy signal. Fine-tune on purchases or add-to-cart
    events, weighting by signal strength.
  • One embedding set for search and recommendations. These are different tasks: search optimises
    relevance to a query, recommendations optimise engagement and conversion.
  • No retraining. The catalogue changes and behaviour moves. Embeddings trained six months ago do
    not reflect seasonal trends or new arrivals.