How the attention mechanism works
Attention operates on three matrices computed from the input vectors: Query (Q), Key (K) and
Value (V). For each token, the similarity between its Query and the Keys of every other token is
computed — those are the attention weights. A weighted sum of the Values gives the token a
context-enriched representation.
The scaled dot-product attention formula:
Attention(Q, K, V) = softmax(Q * K^T / sqrt(d_k)) * V
Dividing by the square root of d_k prevents vanishing gradients at high dimensionality.
Why it matters
Before attention, neural networks processed sequences step by step, each token passing information
to the next. On longer texts the context blurred: by the end of a sentence the model barely
remembered its beginning.
Attention let every token see every other token in the sequence directly. That brought three
fundamental improvements:
- Long-range dependencies — the model weighs tokens at the start and the end of a sequence
equally well. - Parallelism — the whole weight matrix is computed in one step, which trains far more
efficiently on a GPU than an RNN does. - Interpretability — attention weights can be visualised, showing what the model is looking at
while processing a given token.
Multi-head attention
| Parameter | Value |
|---|---|
| Number of heads (standard) | 8–16 |
| What each head learns | Different aspects: syntax, semantics, coreference |
| Output | Concatenation of the heads, then a linear projection |
Heads specialise: one tracks syntactic dependencies, another semantic similarity, a third positional
patterns.
Use in recommendations and search
In recommender systems, attention is used to model session behaviour: the model weighs which of the
previously viewed products are most relevant for predicting the next one. Unlike a simple average of
embeddings, attention takes the order and the context of interactions into account.
Important: attention is a computationally expensive operation. Inference on transformer models
takes significant resources, which affects the latency of real-time recommendation APIs.