The idea: separate the query from the catalogue
The point of a two-tower model is to keep the search for suitable items fast at any catalogue size.
To achieve that, user and item are encoded separately into one and the same vector space.
Relevance is measured by the dot product, or cosine similarity, of the two vectors.
The key property: item vectors are computed once and stored in a vector database. An online
request only needs the user vector plus an ANN lookup, which takes milliseconds regardless of
catalogue size.
query_vec = user_tower(user_features)
candidate_vecs = item_tower(item_features) ← precomputed
scores = dot_product(query_vec, candidate_vecs)
top_k = ANN_search(query_vec, index)
Architecture and training
Each tower is a neural network of arbitrary architecture. The standard option is several fully
connected layers (an MLP) over the input features. For the user’s interaction history, mean pooling
of embeddings or a transformer is used.
The model trains on interaction pairs. The most widespread approach is in-batch negatives: for
every positive example — a user and a purchased item — all the other items in the batch serve as
negatives. That is cheap and effective enough once the batch is large.
| Component | E-commerce example |
|---|---|
| User tower input | The last N items, categories of interest, RFM features |
| Item tower input | Item ID, category, text embedding of the title, price, brand |
| Loss function | In-batch softmax, BPR (Bayesian Personalized Ranking) |
| Quality metric | Recall@K, NDCG@K |
A two-stage architecture: retrieval plus ranking
A two-tower model usually serves the first stage — candidate retrieval: pulling a few thousand
candidates out of a catalogue of millions in milliseconds. At the second stage a heavier model,
gradient boosting or a reranking network, orders those candidates using additional features.
That split makes it possible to apply expensive models to a small candidate set and avoid scoring
the whole catalogue on every request.
Tip: on item cold start a two-tower model beats matrix factorization — the item tower uses
text attributes that need no interaction history at all. For a new user the same logic applies
through contextual features.