How the attention mechanism works

Attention operates on three matrices computed from the input vectors: Query (Q), Key (K) and
Value (V). For each token, the similarity between its Query and the Keys of every other token is
computed — those are the attention weights. A weighted sum of the Values gives the token a
context-enriched representation.

The scaled dot-product attention formula:

Attention(Q, K, V) = softmax(Q * K^T / sqrt(d_k)) * V

Dividing by the square root of d_k prevents vanishing gradients at high dimensionality.

Why it matters

Before attention, neural networks processed sequences step by step, each token passing information
to the next. On longer texts the context blurred: by the end of a sentence the model barely
remembered its beginning.

Attention let every token see every other token in the sequence directly. That brought three
fundamental improvements:

  • Long-range dependencies — the model weighs tokens at the start and the end of a sequence
    equally well.
  • Parallelism — the whole weight matrix is computed in one step, which trains far more
    efficiently on a GPU than an RNN does.
  • Interpretability — attention weights can be visualised, showing what the model is looking at
    while processing a given token.

Multi-head attention

Parameter Value
Number of heads (standard) 8–16
What each head learns Different aspects: syntax, semantics, coreference
Output Concatenation of the heads, then a linear projection

Heads specialise: one tracks syntactic dependencies, another semantic similarity, a third positional
patterns.

Use in recommendations and search

In recommender systems, attention is used to model session behaviour: the model weighs which of the
previously viewed products are most relevant for predicting the next one. Unlike a simple average of
embeddings, attention takes the order and the context of interactions into account.

Important: attention is a computationally expensive operation. Inference on transformer models
takes significant resources, which affects the latency of real-time recommendation APIs.