Where the transformer came from
Before 2017 recurrent neural networks (RNN, LSTM) dominated: they processed text sequentially, word
after word. That created two limitations — slow training, because nothing could be parallelised, and
the tendency to forget context from the beginning of a long text.
The paper “Attention is All You Need” (Vaswani et al., 2017) proposed dropping recurrence
altogether: take the attention mechanism alone, apply it to every position at once, and add
positional encodings to preserve token order.
How self-attention works
For every token the model computes three vectors: Query (Q), Key (K) and Value (V). The
attention of token A to token B is the dot product of Q_A and K_B, normalised by softmax. The result
is a weighted sum of the V vectors of all tokens. Put another way: each token votes on which other
tokens matter for interpreting it.
Attention(Q, K, V) = softmax(QK^T / √d_k) · V
Multi-head attention runs this process in parallel across several projections and concatenates the
results, which lets the model capture different types of dependency at the same time.
Encoders and decoders in e-commerce
| Type | Examples | E-commerce use |
|---|---|---|
| Encoder-only | BERT, RoBERTa | Semantic search, query-product matching |
| Decoder-only | GPT, Llama | AI Shopping Assistant, description generation |
| Encoder-decoder | T5, BART | Review summarisation, multilingual search |
Search tasks usually call for encoders: they turn the shopper’s query and the product description
into vectors that can be compared by cosine similarity. Conversational scenarios call for decoders
or encoder-decoder models.
Typical implementation mistakes
- Overestimating the model size. For recommendation tasks BERT-base (110M parameters) is often
enough — heavy models of 7B parameters and above do raise quality, but they lose on inference
latency. - Ignoring context length. Most transformers cap the input sequence length (512 or 2,048
tokens). Long product descriptions or user histories need truncation or a specialised
architecture. - Fine-tuning without domain data. A general BERT understands category-specific terminology
poorly. Fine-tuning on product descriptions and real search queries improves quality
substantially.