Jev yes TypeSafe AI A brand-new AI model launched this week, founded by Diogo Almeida— he was formerly a researcher at OpenAI and helped build ChatGPT, and co-invented RLHF (Reinforcement Learning from Human Feedback), the foundational technology underpinning today’s large model era. What makes Jev uniquely special is:It is not a large language model (LLM), it does not output text—it outputs probabilities.
core competencies
After leaving OpenAI, Almeida founded TypeSafe AI with a singular mission: to make AI truly “useful,” not merely fluent in human language. He observed that, over the past four years, the industry has optimized for “human language,” yet computers actually require a different language—Structured decisions. Jev was built for exactly this purpose.
- Outputs probabilities instead of text: Jev outputs “calibrated decisions”—judgments accompanied by confidence scores—not blocks of text
- No hallucinations: Because the output space is predefined by the user, the model cannot “fabricate” non-existent answers
- Extremely low cost: Output tokens are completely free; input tokens are billed per billion—not per million, as is standard across the industry
- Extremely fast: By eliminating the language generation step, inference speed increases dramatically
- “System One” model: Focuses on intuitive judgment rather than step-by-step reasoning
Technically, Jev is trained exclusively on synthetic data using a method Almeida calls “Reinforcement Learning from Calibrated Decisions” (RLCD). The company keeps its architecture confidential, and external speculation suggests it may be built atop an open-weight LLM.
User experience/limitations
Jev is currently offered via API, primarily targeting developers for tasks such asclassification, routing, and security validation—automated tasks of this kind. At launch, demand surged so high that the company’s API briefly became unavailable to users.
Real-world test results are impressive: Vercel engineers replaced OpenAI’s ChatGPT Luna 5.6 for command safety classification and achieved a speedup of 5x to 18xwhile also improving accuracy; Bryo AI’s CTO benchmarked it against Google Gemini for email classification—Gemini delivered marginally higher accuracy, but at a cost 10x to 20x higher.
Its limitations are clear-cut:Jev cannot generate textand is suited only for “making judgments,” not “writing content”; additionally, it partially shifts the “hallucination” problem to users—when the model returns 50% confidence, you must decide whether to trust it. It functions more like an “intelligent validator” or “low-cost router” for LLMs than a replacement.
Overall Score
| 维度 | Score | evaluate |
|---|---|---|
| functional completeness | 7.5 / 10 | Specialized in decision-making/classification—not a general-purpose model—with narrow capabilities |
| 易用性 | 7.0 / 10 | Developer-focused, requiring predefined output structure |
| Cost-effectiveness | 9.5 / 10 | Output is free; input is billed per billion tokens—extremely low cost |
| 中文支持 | 8.0 / 10 | Language-agnostic, fully applicable to Chinese-language scenarios |
| 输出质量 | 8.0 / 10 | Calibrated probabilities are reliable and hallucination-free, though slightly outperformed by LLMs on some tasks |
Overall rating: 8.0/10
Jev is a “counter-trend” product: at a time when large models grow ever larger and more expensive, it trades away language generation to gain speed, cost efficiency, and reliability. For developers needing massive-scale automated judgments, it may be a more pragmatic choice than LLMs. To discover more useful AI tools, visit AI Dash.

Pingback: Decision models surge: Amazon open-sources Strands Decider 2B, a locally-runnable AI routing tool — AI Dash