How a context window works

The context window is the model’s “working memory” for a single request. Everything the model has to consider is sent to it again on every call and measured in tokens — the word fragments that a tokenizer splits text into. The window holds:

  • system instructions — the assistant’s role, tone and restrictions;
  • descriptions of the tools the model can call;
  • the conversation history;
  • data retrieved by searching the catalogue;
  • the shopper’s current question;
  • the model’s answer, including reasoning tokens if there are any.

Window size: orders of magnitude

Date What is known Source
February 2024 Gemini 1.5 Pro: standard window of 128,000 tokens, up to 1 million in preview; Gemini 1.0 had 32,000 Google blog
September 2026 Current Claude models — 1 million tokens; some earlier ones (for example, Claude Sonnet 4.5) — 200,000 Anthropic documentation

The figures change with every model generation, so an assistant’s architecture should not depend on a specific window size.

Why a large window does not solve the problem

  • Cost. Input tokens are billed on every request. The longer the context, the more each turn of the conversation costs.
  • Latency. A long input takes longer to process at the inference stage, while the shopper waits for a reply in the chat.
  • Lost in the middle. In the “Lost in the Middle” paper (Liu et al., 2023), models used information from the beginning and end of the input best and noticeably worse from the middle. Anthropic calls the general loss of accuracy as context grows “context rot”.

Hence the standard approach: rather than enlarging the window, select what goes into it through RAG. Deciding what goes into the window and in what order is the job of context engineering.

Example: a shopping assistant’s context budget

The numbers below are illustrative — real usage is measured with the chosen model’s tokenizer.

What goes into the window Tokens (estimate)
System instructions and merchandising rules 1,500
Descriptions of three tools: search, basket, order status 1,000
Conversation history, 8 turns 1,500
20 product records at 250 tokens each 5,000
Model answer 500
Total 9,500

A request like this fits comfortably into the window of any current model. The bottleneck is not the window size but which 20 products made it in and how much each turn costs.

Checklist:

  • pass only the fields you need in a product record — title, price, availability, 3–5 key attributes — not the full HTML description;
  • put the rules and the most relevant products at the start, and the current question at the end;
  • compress old turns into a summary, and keep long-term facts about the shopper outside the window — see agent memory;
  • track tokens per turn alongside latency and conversion: they are a direct cost line.