Where the transformer came from

Before 2017 recurrent neural networks (RNN, LSTM) dominated: they processed text sequentially, word
after word. That created two limitations — slow training, because nothing could be parallelised, and
the tendency to forget context from the beginning of a long text.

The paper “Attention is All You Need” (Vaswani et al., 2017) proposed dropping recurrence
altogether: take the attention mechanism alone, apply it to every position at once, and add
positional encodings to preserve token order.

How self-attention works

For every token the model computes three vectors: Query (Q), Key (K) and Value (V). The
attention of token A to token B is the dot product of Q_A and K_B, normalised by softmax. The result
is a weighted sum of the V vectors of all tokens. Put another way: each token votes on which other
tokens matter for interpreting it.

Attention(Q, K, V) = softmax(QK^T / √d_k) · V

Multi-head attention runs this process in parallel across several projections and concatenates the
results, which lets the model capture different types of dependency at the same time.

Encoders and decoders in e-commerce

Type Examples E-commerce use
Encoder-only BERT, RoBERTa Semantic search, query-product matching
Decoder-only GPT, Llama AI Shopping Assistant, description generation
Encoder-decoder T5, BART Review summarisation, multilingual search

Search tasks usually call for encoders: they turn the shopper’s query and the product description
into vectors that can be compared by cosine similarity. Conversational scenarios call for decoders
or encoder-decoder models.

Typical implementation mistakes

  • Overestimating the model size. For recommendation tasks BERT-base (110M parameters) is often
    enough — heavy models of 7B parameters and above do raise quality, but they lose on inference
    latency.
  • Ignoring context length. Most transformers cap the input sequence length (512 or 2,048
    tokens). Long product descriptions or user histories need truncation or a specialised
    architecture.
  • Fine-tuning without domain data. A general BERT understands category-specific terminology
    poorly. Fine-tuning on product descriptions and real search queries improves quality
    substantially.