How a reasoning model works
A standard LLM generates the answer token by token as soon as it receives the request. A reasoning model first writes a long internal chain of reasoning: it breaks the task into steps, tries options, spots mistakes and backtracks. The user sees only the final answer — in OpenAI o1 the chain of reasoning is hidden by design.
Models learn this behaviour through reinforcement learning. In the DeepSeek-R1 paper, the authors showed that reasoning ability can be developed through “pure” reinforcement learning, without human-labelled reasoning examples. The price is paid in computation at the inference stage: the longer the model “thinks”, the more accurate the answer and the more expensive it is.
Milestones
| Date | Event |
|---|---|
| 12 September 2024 | OpenAI releases o1-preview and o1-mini; in the announcement the company reports that o1 solves 83% of problems on the AIME maths exam versus 13% for GPT-4o |
| 5 December 2024 | The full version of o1 is released |
| 20 January 2025 | DeepSeek-R1: open weights, MIT licence, six smaller distilled models |
| 24 February 2025 | Claude 3.7 Sonnet — a hybrid model: standard mode and extended thinking in one model |
As of September 2026, the depth of reasoning in major APIs is set by a request parameter — reasoning effort at OpenAI, a budget or adaptive thinking mode at Anthropic.
Pros and cons for e-commerce
| Assistant task | Standard LLM | Reasoning model |
|---|---|---|
| Delivery, returns, order status | Fast and good enough | Overkill: slower and more expensive |
| Selection by one or two parameters | Copes well | Small gain |
| Comparing 5 products on 6 criteria | May mix up specifications | A strength |
| Bundle within a budget, compatibility check | More often gets arithmetic and conditions wrong | A strength |
| Multi-step agent actions | More often loses the plan | Plans and checks each step |
Reasoning does not eliminate hallucination: if the context contains no data about a product, the model can still invent it — just with better arguments.
In practice: routing in an online assistant
- Classify the request. A fast classifier or a set of rules decides whether the question is simple or complex.
- Simple questions go to a fast model. FAQs, availability, delivery, size checks.
- Complex ones go to a reasoning model. Comparisons, bundles, compatibility, multi-step AI agent scenarios.
- Set a budget. Cap the depth of reasoning and the response time; for the first reaction, show a short line such as “Comparing the models against your criteria”.
- Measure. Compare p95 latency, cost per conversation and conversion in an A/B test, not just answer quality on test questions.