What Jev is
Jev is the first System One model from TypeSafe AI, an AI lab in San Francisco. Early access opened on 25 September 2026. The company’s CEO is Diogo Almeida; according to TypeSafe AI, he helped create RLHF and InstructGPT, the methods behind ChatGPT.
The model is built not for conversation but for decisions inside software: classifying, scoring and choosing a branch of a workflow where hand-written if-else rules are too brittle and a full LLM is too slow and expensive. It is called through an HTTP API (POST /v1/systemone) or SDKs for Python and JavaScript.
How it works
A request contains a state — text or JSON with the context — and a set of questions. All questions are evaluated in parallel and independently against the same state.
| Question type | What it asks | Store example | What it returns |
|---|---|---|---|
| Choice | Which option from a list | Who should handle the request: billing, delivery or returns | The chosen option, a probability for each, confidence |
| Score | A rating on a described scale | How frustrated the customer is: 0 is calm, 2 is very frustrated | A score, probabilities for each level, confidence |
| Noul | Whether a statement is true | Is the customer asking for a refund | The probability of yes, from 0 to 1 |
The model is trained with RLCD (Reinforcement Learning for Calibrated Decisions): its probabilities are optimised to reflect real uncertainty. Jev is not fine-tuned on customer data — every account uses the same weights. You adapt it to your task through the request: reference data goes in the state, rules and edge cases in the question instructions and criteria.
Specifications and company claims
| Parameter | Value as of September 2026 |
|---|---|
| Version | jev-1.13.0 (alias jev-latest) |
| Context | 64,000 tokens per request; 32,000 for the state plus the longest question |
| Input | Text only: a string, JSON or an array of strings |
| Price | $42 per billion input tokens; output is free |
| Limits | 250,000 tokens per second, 1,200 requests per minute |
By TypeSafe AI’s measurements, on System One tasks Jev is 193.6 times faster and 444.6 times cheaper than LLMs: an answer arrives in 70–500 milliseconds against 3–329 seconds. The reference in that comparison was the predictions of large models — GPT-6 Astra and Fable 5.1. These are the vendor’s own figures; as of September 2026 there are no independent benchmarks.
Where it could help in e-commerce
TypeSafe AI’s documentation lists these scenarios for marketplaces and support: normalising seller listings, extracting attributes from titles, detecting prohibited items and fake reviews, and classifying and routing requests by intent. A separate scenario is scoring catalogue passages before they reach a shopping assistant’s answer. As of September 2026 no retailer case studies have been published.
Limitations and how to test
The company openly publishes the weak spots of version 1.13: literal reading of wording, counting and arithmetic, date comparison, indirect references, a large state full of irrelevant detail, contradictory instructions and adversarial content such as prompt injection.
- Test non-English content separately: English is the primary training language.
- Build a reference set and compare Jev with your current solution on accuracy, latency and price — the same method as LLM evaluation.
- Set confidence thresholds: above the threshold, act automatically; below it, send the case to an agent or a reasoning model.
- Once thresholds are tuned, pin
jev-1.13.0instead ofjev-latest: the alias moves to a new model without warning.