The most talked-about AI model of September can’t write a single sentence. Jev, released September 15 by the startup TypeSafe AI, doesn’t generate text at all. Instead, it reads your input and returns a decision: yes or no, a pick from your list, a score — each with a calibrated probability attached.

TypeSafe calls this new category “System One” models (after Kahneman’s Thinking, Fast and Slow; “Jev” itself is named after economist William Stanley Jevons). Independent commentators like Simon Willison suggest the clearer name is simply decision models.

How it works

A normal LLM takes your prompt and generates tokens one by one; you parse whatever comes out and hope the JSON is valid. Jev flips the interface:

  1. You send a state (text, JSON, or an array of strings).
  2. You define typed questions in advance.
  3. Jev evaluates all questions in parallel, in a single pass, and returns typed answers with probability distributions — nothing to parse, schema-conformance guaranteed by construction.

Three primitives: Choice (pick from up to 255 options), Score (2–10 ordered levels), and Noul (yes/no probability 0–1; “Noul” is short for Bernoulli). The training method is called RLCD — Reinforcement Learning for Calibrated Decisions — which replaces RLHF with training for calibrated uncertainty, so stated confidence actually tracks real accuracy.

The pitch, in TypeSafe’s words: “a frontier-intelligence function call — unstructured state in, typed probabilistic decisions out.” Vendor-claimed specs: $0.042 per million input tokens with output tokens free, 70–500 ms latency, 64K context. Note the “can’t hallucinate” claim means format-safe, not error-free: independent tests found wrong answers delivered at 80%+ confidence on math problems.

The company: co-founder/CEO Diogo Almeida, an ex-OpenAI researcher (RLHF/InstructGPT/ChatGPT) who left in 2024, emerged from stealth with the launch, backed by a $40M seed led by DCVC. The Information later reported talks of a $1B+ raise at a $10B+ valuation — unconfirmed, and Almeida declined to comment.

Real-world use cases

The strongest third-party evidence comes from Vercel: Jev was the fastest-adopted model in AI Gateway history — roughly 13% of paid teams within 24 hours, 2× the GPT-5.6 family. One caveat: that was measured during a free-access window (free through Sept 25), so it’s trial velocity, not paid conversion. Cloudflare, LangChain, and Langfuse integrated within three days of launch.

What people actually use it for:

  • Agent infrastructure at Vercel: tool/subagent selection, continue/retry/stop decisions, urgency and risk scoring, output verification, guardrails — and escalating uncertain cases to humans.
  • Content moderation: in an independent 900-example benchmark (AI/ML API, Sept 2026), Jev had the best F1 of six models on offensive-content moderation, beating Claude Opus 5.5 and GPT-6 Sol — while costing 84–150× less than Opus 5.5.
  • The filter pattern: as a first-pass filter in front of Opus 5.5, Jev matched Opus’s accuracy on classification tasks at 38–45% of the cost. This looks like the honest sweet spot: fast, cheap, well-calibrated triage in front of an LLM — not an LLM replacement.
  • Routing and scoring: ticket routing, lead scoring, review routing, fraud signals.

Be skeptical of the headline numbers, though. Almeida told the WSJ that ~25% of the Fortune 500 “use” Jev with traffic around a trillion tokens a day — but no companies are named, “in use” is undefined, and nothing about revenue or retention is disclosed. Treat it as a vendor claim until proven otherwise. Also worth knowing: prompt injection can sway Jev’s verdicts, so wiring it into agent decision loops is running ahead of governance (VentureBeat).

The clone wars

Here’s the remarkable part: within 16 days of Jev’s launch, the copycats arrived in force.

Date Clone Backstory
Sept 29 OpenAI Decisions API (DevDay) Gives Luna a predefined set of options to choose between. Limited preview, no public docs or pricing at launch. Almeida joked on X about the “clone wars”; TypeSafe declined to comment.
Sept 30 Databricks ai_decide Native AI Function via SQL + REST, Jev-API compatible. For model routing, review tagging, agent eval.
Oct 1 Cloudflare Clef / Clef-flash The most aggressive: the first models Cloudflare ever trained itself (27B + 9B, Qwen-based), Apache 2.0 open weights, Jev-API compatible — plus a vision encoder and 64K context Jev lacks. Cloudflare claims it beats Jev on 3 of 4 of TypeSafe’s own evals with ~13× lower latency. All vendor-run and unreproduced; one independent spot-check found the opposite on nuanced triage.
Oct 1 AWS Strands Decider 2B The best story: it started as distinguished engineer Marc Brooker’s side project after he saw Jev — he built a homebrew version (~2B params, Qwen3.5 + LoRA + a tiny “pointer head”), briefly topped JevBench for its size class, and AWS productized it in about two weeks. Fully open source.

On top of that, community open weights appeared within ~10 days: Laya, Kev, Von, Tev1-4B, OpenJev, GLiNER2.5-Decide — trailing Jev zero-shot but competitive after fine-tuning.

We’ve seen this movie before, just never this fast: ChatGPT (Nov 2022) → Google rushed Bard to testers ~2.5 months later; DeepSeek R1 (Jan 2025) → distilled clones within days and a $589B single-day Nvidia wipeout. Jev compresses the same pattern into days.

Good reads