How a reasoning model works

A standard LLM generates the answer token by token as soon as it receives the request. A reasoning model first writes a long internal chain of reasoning: it breaks the task into steps, tries options, spots mistakes and backtracks. The user sees only the final answer — in OpenAI o1 the chain of reasoning is hidden by design.

Models learn this behaviour through reinforcement learning. In the DeepSeek-R1 paper, the authors showed that reasoning ability can be developed through “pure” reinforcement learning, without human-labelled reasoning examples. The price is paid in computation at the inference stage: the longer the model “thinks”, the more accurate the answer and the more expensive it is.

Milestones

Date Event
12 September 2024 OpenAI releases o1-preview and o1-mini; in the announcement the company reports that o1 solves 83% of problems on the AIME maths exam versus 13% for GPT-4o
5 December 2024 The full version of o1 is released
20 January 2025 DeepSeek-R1: open weights, MIT licence, six smaller distilled models
24 February 2025 Claude 3.7 Sonnet — a hybrid model: standard mode and extended thinking in one model

As of September 2026, the depth of reasoning in major APIs is set by a request parameter — reasoning effort at OpenAI, a budget or adaptive thinking mode at Anthropic.

Pros and cons for e-commerce

Assistant task Standard LLM Reasoning model
Delivery, returns, order status Fast and good enough Overkill: slower and more expensive
Selection by one or two parameters Copes well Small gain
Comparing 5 products on 6 criteria May mix up specifications A strength
Bundle within a budget, compatibility check More often gets arithmetic and conditions wrong A strength
Multi-step agent actions More often loses the plan Plans and checks each step

Reasoning does not eliminate hallucination: if the context contains no data about a product, the model can still invent it — just with better arguments.

In practice: routing in an online assistant

  • Classify the request. A fast classifier or a set of rules decides whether the question is simple or complex.
  • Simple questions go to a fast model. FAQs, availability, delivery, size checks.
  • Complex ones go to a reasoning model. Comparisons, bundles, compatibility, multi-step AI agent scenarios.
  • Set a budget. Cap the depth of reasoning and the response time; for the first reaction, show a short line such as “Comparing the models against your criteria”.
  • Measure. Compare p95 latency, cost per conversation and conversion in an A/B test, not just answer quality on test questions.