What a token is

A language model does not work with letters or words directly. Text is first split into tokens —
sub-word units defined by a tokenization algorithm (BPE, WordPiece, SentencePiece). A token is a
vocabulary unit belonging to that specific model.

Word: "personalization"
Tokens (example): ["personal", "iz", "ation"]  →  3 tokens

Word: "cat"
Tokens: ["cat"]  →  1 token

Inflected languages tokenise less efficiently than English: complex morphology — cases, aspects,
conjugations — produces far more surface forms, and a tokenizer vocabulary trained mostly on English
breaks those words into smaller pieces. The German compound “Personalisierung” splits into four or
five tokens where the English equivalent takes three.

The context window and its limits

The context window is the maximum number of tokens an LLM can process in one request — input and
output together. It is a fundamental limitation of the architecture.

Model Context window
GPT-3.5 16K tokens
GPT-4 128K tokens
Claude 3 200K+ tokens

Even 128K tokens is roughly 100,000 words. For a catalogue of 100,000 SKUs or more, that is not
enough. This is why an e-commerce AI Shopping Assistant does not know the whole catalogue from the
prompt — it uses RAG: vector search across the catalogue, retrieval of the relevant products, and
injection of only those into the request context.

Practical consequences for e-commerce AI

Cost optimisation. When an LLM is integrated into e-commerce — an AI Shopping Assistant, product
description generation — the cost equals the number of tokens multiplied by the price per 1K.
Compact prompts and structured data instead of long prose reduce the bill.

Context limits. You cannot pass an entire purchase history and an entire catalogue in one
prompt. RAG solves the problem: search first, then feed only the relevant subset to the LLM.

Non-English text costs more. Tokenizing a morphologically rich language produces roughly 1.5 to
2 times more tokens than English text of the same meaning. That has to be factored into the cost
estimate for any localised AI feature.