TypeSafe AI's "System One" Model: JEV
Released in September 2026 by TypeSafe AI (founded by former OpenAI researcher Diogo Almeida), Jev is a frontier-class AI model that completely abandons the traditional chatbot or generative text paradigm.
Here is a full breakdown of how Jev works, how it diverges from the rest of the AI industry, and the trade-offs involved in its design.
1. Taking Research in the "Opposite Direction"
For years, AI research has been moving in one primary direction: making Large Language Models (LLMs) better at conversational generation, step-by-step reasoning ("System 2" thinking), and open-ended text creation.
TypeSafe AI took Jev in the exact opposite direction. They built a "System One" model (named after the psychological concept of fast, intuitive, automatic thinking).
- No Prose, Only Decisions: Jev does not generate conversational text, summarize documents, or write code.
- Structured Output Only: It takes unstructured input (like an email or a system log) alongside a set of predefined questions and returns only structured, typed data (JSON).
- Guaranteed Schema: Because the developer defines the exact answer space in advance (e.g., asking the model to pick from 3 specific categories), Jev cannot hallucinate outside of those bounds or break the application's code by returning a malformed string.
2. What's the Difference? (Jev vs. Standard LLMs)
While an LLM acts as a generator, Jev acts as a decision layer for software.
| Feature | Generative LLMs (GPT, Claude, etc.) | Jev (TypeSafe AI) |
|---|---|---|
| Primary Output | Free-form text, prose, explanations, code. | Typed decisions with probability & confidence scores. |
| Answer Space | Open-ended and dependent on prompting. | Strictly defined by the developer in advance. |
| Generation Method | Autoregressive (predicts one token at a time). | Parallel Sampling (evaluates all outputs at once). |
| Best Use Cases | Writing, summarizing, brainstorming, coding. | Classification, routing, moderation, risk scoring. |
Jev allows developers to ask three specific types of "primitives" (questions):
- Choice: Select one option from a defined set (up to 255 options).
- Score: Rate the input on a defined ordered scale (e.g., 1 to 10).
- Noul: A strict yes/no boolean judgment.
3. The "New RLHF": RLCD
Most modern chatbots are trained using RLHF (Reinforcement Learning from Human Feedback), which optimizes the model to produce answers that humans find pleasing, helpful, and plausible.
Instead of RLHF, Jev introduces RLCD (Reinforcement Learning for Calibrated Decisions).
- Honest Probabilities over Plausibility: RLCD optimizes for outcome accuracy rather than human preference.
- Self-Aware Confidence: The training ensures that when Jev reports a "high confidence" score alongside its decision, it is genuinely highly accurate. If a decision is borderline, Jev outputs a low confidence score, allowing the software to automatically route the task to a human for manual review.
4. How Jev is So Fast and Cheap
Jev is reportedly 40 to 200 times faster and 40 to 400 times cheaper than standard frontier LLMs. It achieves this primarily through its architecture:
- No Autoregressive Generation: Standard LLMs spend massive amounts of compute generating text one word (token) at a time. Jev skips this entirely.
- Parallel Sampling: Jev evaluates the input state and all requested questions simultaneously in a single parallel pass, returning the structured JSON decision almost instantly (70 to 500 milliseconds).
- Unique Pricing Model: Because Jev doesn't generate long strings of text, TypeSafe AI prices it radically differently. Input tokens cost roughly $42 per billion, and output tokens are entirely free.
5. Drawbacks and Limitations
Because Jev is highly specialized, it comes with strict limitations:
- Cannot Generate Content: Jev cannot write an email, draft a summary, explain its reasoning, or converse with a user.
- No Complex Reasoning Chains: It lacks the step-by-step reasoning capabilities of standard models. It makes split-second "gut" decisions based on its training.
- Small Context Window: Jev is capped at a 64,000-token context window, which is significantly smaller than the 1-million+ token context windows common in modern LLMs. You cannot feed it massive books or vast codebase histories.
- Can Still Be Wrong: While Jev won't hallucinate an answer outside of your predefined categories (meaning it won't break your code's schema), it can still confidently select the wrong category from the list you provided.
Conclusion: Jev is not a replacement for models like GPT or Claude; it is a companion tool. It is designed to sit upstream of generative models to handle high-volume, low-latency classification, routing, and moderation tasks where speed, cost, and strict data formatting are paramount.
Join the conversation