What a token is
A language model does not work with letters or words directly. Text is first split into tokens —
sub-word units defined by a tokenization algorithm (BPE, WordPiece, SentencePiece). A token is a
vocabulary unit belonging to that specific model.
Word: "personalization"
Tokens (example): ["personal", "iz", "ation"] → 3 tokens
Word: "cat"
Tokens: ["cat"] → 1 token
Inflected languages tokenise less efficiently than English: complex morphology — cases, aspects,
conjugations — produces far more surface forms, and a tokenizer vocabulary trained mostly on English
breaks those words into smaller pieces. The German compound “Personalisierung” splits into four or
five tokens where the English equivalent takes three.
The context window and its limits
The context window is the maximum number of tokens an LLM can process in one request — input and
output together. It is a fundamental limitation of the architecture.
| Model | Context window |
|---|---|
| GPT-3.5 | 16K tokens |
| GPT-4 | 128K tokens |
| Claude 3 | 200K+ tokens |
Even 128K tokens is roughly 100,000 words. For a catalogue of 100,000 SKUs or more, that is not
enough. This is why an e-commerce AI Shopping Assistant does not know the whole catalogue from the
prompt — it uses RAG: vector search across the catalogue, retrieval of the relevant products, and
injection of only those into the request context.
Practical consequences for e-commerce AI
Cost optimisation. When an LLM is integrated into e-commerce — an AI Shopping Assistant, product
description generation — the cost equals the number of tokens multiplied by the price per 1K.
Compact prompts and structured data instead of long prose reduce the bill.
Context limits. You cannot pass an entire purchase history and an entire catalogue in one
prompt. RAG solves the problem: search first, then feed only the relevant subset to the LLM.
Non-English text costs more. Tokenizing a morphologically rich language produces roughly 1.5 to
2 times more tokens than English text of the same meaning. That has to be factored into the cost
estimate for any localised AI feature.