How Item2Vec works

The method borrows the Word2Vec architecture from NLP. In the text setting, Word2Vec learns which
words occur near each other inside sentences. Item2Vec does the same thing, except the sentences are
shopper sessions — sequences of viewed or purchased products — and the words are product IDs.

During training, every product gets a vector of fixed dimensionality, typically 64 to 256 numbers.
Two products that regularly appear in the same sessions receive close vectors. Afterwards the
distance in vector space reads directly: the smaller the cosine distance between two products, the
more likely they are to be bought together.

Product A: [0.12, -0.45, 0.87, ..., 0.33]  <- a jacket
Product B: [0.14, -0.41, 0.85, ..., 0.31]  <- a similar jacket
Product C: [-0.72, 0.90, -0.11, ..., 0.05] <- an umbrella (different cluster)

cosine_similarity(A, B) = 0.97  <- high similarity
cosine_similarity(A, C) = 0.22  <- low similarity

Advantages over classical methods

Criterion Matrix factorization Item2Vec
User identification Required Not required
Anonymous sessions Awkward Native
Scaling to a large catalogue Moderate High
Order of items within a session Ignored in the basic form Available through skip-gram
Adding new products Recompute the matrix Incremental further training

Important: Item2Vec captures joint consumption, not product attributes. Two physically
dissimilar products — a sports bottle and a set of dumbbells — can end up close in vector space if
they are regularly bought together, and behaviourally that is correct.

Where it is used in e-commerce

Similar items recommendations. For a given product, the k nearest neighbours by vector are
taken. An ANN query to a vector database runs in milliseconds even across catalogues of millions.

Cross-sell in the basket. The items already in the basket are averaged into a centroid vector,
and the nearest neighbours not yet in the basket are retrieved.

Session-based recommendations. Averaging the vectors of the products viewed in the current
session produces a session vector; querying the database with it returns relevant next products with
no user history at all.

Limitations

Cold start for products. A new item added to the catalogue after the last training run has no
vector until the next cycle. The usual workaround is to assign it the vector of its nearest
neighbour by attributes.

Session length. Very short sessions of one or two products give little context. The model is
more reliable for products with moderate view frequency — the bestsellers and the deep tail of the
catalogue are both represented worse.

No explicit user context. Item2Vec knows nothing about demographics, purchase history or the
loyalty of a specific shopper. For personalized recommendations it is combined with user embeddings
or with a two-tower architecture.