How a context window works
The context window is the model’s “working memory” for a single request. Everything the model has to consider is sent to it again on every call and measured in tokens — the word fragments that a tokenizer splits text into. The window holds:
- system instructions — the assistant’s role, tone and restrictions;
- descriptions of the tools the model can call;
- the conversation history;
- data retrieved by searching the catalogue;
- the shopper’s current question;
- the model’s answer, including reasoning tokens if there are any.
Window size: orders of magnitude
| Date | What is known | Source |
|---|---|---|
| February 2024 | Gemini 1.5 Pro: standard window of 128,000 tokens, up to 1 million in preview; Gemini 1.0 had 32,000 | Google blog |
| September 2026 | Current Claude models — 1 million tokens; some earlier ones (for example, Claude Sonnet 4.5) — 200,000 | Anthropic documentation |
The figures change with every model generation, so an assistant’s architecture should not depend on a specific window size.
Why a large window does not solve the problem
- Cost. Input tokens are billed on every request. The longer the context, the more each turn of the conversation costs.
- Latency. A long input takes longer to process at the inference stage, while the shopper waits for a reply in the chat.
- Lost in the middle. In the “Lost in the Middle” paper (Liu et al., 2023), models used information from the beginning and end of the input best and noticeably worse from the middle. Anthropic calls the general loss of accuracy as context grows “context rot”.
Hence the standard approach: rather than enlarging the window, select what goes into it through RAG. Deciding what goes into the window and in what order is the job of context engineering.
Example: a shopping assistant’s context budget
The numbers below are illustrative — real usage is measured with the chosen model’s tokenizer.
| What goes into the window | Tokens (estimate) |
|---|---|
| System instructions and merchandising rules | 1,500 |
| Descriptions of three tools: search, basket, order status | 1,000 |
| Conversation history, 8 turns | 1,500 |
| 20 product records at 250 tokens each | 5,000 |
| Model answer | 500 |
| Total | 9,500 |
A request like this fits comfortably into the window of any current model. The bottleneck is not the window size but which 20 products made it in and how much each turn costs.
Checklist:
- pass only the fields you need in a product record — title, price, availability, 3–5 key attributes — not the full HTML description;
- put the rules and the most relevant products at the start, and the current question at the end;
- compress old turns into a summary, and keep long-term facts about the shopper outside the window — see agent memory;
- track tokens per turn alongside latency and conversion: they are a direct cost line.