Where the name comes from

In Thinking, Fast and Slow, Daniel Kahneman described two systems of thought: System 1 is fast and intuitive, System 2 is slow and deliberate. In AI terms, System 2 corresponds to reasoning models, which build a long chain of reasoning before they answer. A System One model sits at the opposite pole: a fast, narrow judgment that a program receives in a fraction of a second and uses without parsing any text.

The term was introduced by TypeSafe AI, a San Francisco AI lab. On 25 September 2026 it opened early access to the first model of this class, Jev. As of September 2026 no other models carry the name, so the concept currently describes one company’s approach.

How it differs from an LLM

An LLM is trained to generate text for a person. A System One model also understands natural language, but it answers in a format that code can act on directly.

LLM System One model
Output Text: a reply, code, an explanation An option from a defined set, a score on a scale or the probability of yes
How the format is set By a prompt instruction the model may drift from By the question schema, which the model cannot leave
Uncertainty Usually not reported, or overstated A probability for each option plus a separate confidence estimate
Who the answer is for A person A program deciding what to do next

Jev offers three question types. Choice picks one option from a list, for example billing, delivery or returns. Score rates content on a described scale, for example how frustrated a customer is. Noul answers a yes/no question with a probability: is the customer asking for a refund.

Calibrated confidence

Calibration means that probability matches observed frequency: among answers given with 90% confidence, about 90% should be correct. It is a property of a group of answers, not a guarantee for any single decision. The practical value is that confidence becomes a second axis: the answer says what to do, and confidence says whether to do it automatically or pass the case to an agent or a stronger model. That puts a human in the loop only where the model is unsure.

Important: typed output removes part of the hallucination risk — the model will not invent a category that does not exist. It can still choose the wrong one.

E-commerce tasks

TypeSafe AI’s documentation proposes the following scenarios for marketplaces and customer support. Treat them as a class of tasks rather than proven case studies:

  • classifying and normalising product listings from different sellers;
  • extracting attributes from titles and descriptions;
  • detecting prohibited items, counterfeit signals and fake reviews;
  • recognising the intent of a request and routing it to an automated flow, an agent or an LLM;
  • scoring catalogue passages before they reach a shopping assistant’s answer.

How to evaluate before adopting

  • Build a reference set of real requests or listings with correct answers — the same logic as LLM evaluation.
  • Compare accuracy, latency and cost with your current solution: rules, a classic classifier or an LLM.
  • Check calibration on your own data: what share of high-confidence answers turn out wrong.
  • Test non-English content separately: according to the company, Jev works best in English.
  • The speed and price figures as of September 2026 are the vendor’s own measurements; there are no independent comparisons yet.