Back to feed

What Is Jev? ChatGPT Co-Creator's 200x Faster Decision Model Explained

Diego Almeida, co-creator of ChatGPT, spent two years in stealth to launch Jev — a non-generative decision engine. 20–200x faster, up to 400x cheaper at $42 per billion input tokens and free output, this System 1 model classifies everything from email to browser automation in milliseconds.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — 2z-7pIj57f8
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

What Jev Is and Why It's Named After Jevons

Jev is TypeSafe AI's first System 1 model, named for William Stanley Jevons and his paradox: when efficiency rises, consumption rises. Founder Diego Almeida co-invented ChatGPT at Google Brain and OpenAI and says he spent two years in stealth asking why superhuman chat models did not lead to AGI. His answer is a training shift that produces a decision engine that does not generate text.

The architecture has no autoregressive loop. Classic models stream the next token; Jev takes a state and a set of questions and scores them all in parallel. The team claims 20 to 200 times faster and 40 to 400 times cheaper than frontier chat models. Real-time Doom, half-second wiki races and 7-second flight searches are presented as visual proof of that claim.

The framing comes from Daniel Kahneman's Thinking, Fast and Slow. System 1 is fast and intuitive, System 2 slow and deliberate. Frontier models like Claude and GPT behave like System 2; Jev is built to be System 1. The idea is to move a System 2 task into System 1 through practice, like learning to drive and then doing it on autopilot.

RLCD and Three Primitives: Choice, Score, Bool

Jev is trained with RLCD, Reinforcement Learning for Calibrated Decisions, instead of RLHF. Where RLHF optimizes for human preferences and can amplify hallucination, RLCD optimizes for well-calibrated probabilities. TypeSafe advertises zero hallucination for decision tasks because every answer returns a distribution with confidence rather than a freeform sentence.

Interaction relies on three primitives. Choice picks one of 2 to 255 predefined options and returns a distribution, for example fraud versus clean versus human review, or billing versus technical versus sales. Score places a case on a rubric of 2 to 11 levels and returns a value between 0 and 10, such as hobbyist to enterprise or cosmetic to blocking. Bool answers a yes-or-no proposition with a probability between 0 and 1.

Every request follows the same shape: a state on the left, questions on the right. The state can be an email object, an invoice JSON, a game screen description or a code diff. Instructions, choices and optional examples define each question. Because there is no text generation, the output is always type-safe JSON that feeds directly into an if statement, and multiple questions for the same state run in parallel for thousands of decisions per second.

Price, Speed and Limits: What the Numbers Say

Pricing breaks expectations. Input is $42 per billion tokens, which is $0.042 per million. That is about three and a half times cheaper than DeepSeek V4.1 Flash at $0.15. Output tokens are not metered at all; the team says they are too cheap to meter. An early Vercel evaluation noted that a classification workload done with Gemini Flash ran 6 times faster on Jev and saturated the evaluation.

Latency is the headline. Typical playground response is 100 to 200 milliseconds, often around 150. For many users network latency exceeds model latency. Even the fastest LLMs at 250 to 300 tokens per second are an order of magnitude slower when asked to produce the same structured JSON. The context window is 64k input tokens, roughly 6 percent of Astra or Opus at 1 to 1.5 million, so Jev trades breadth for speed.

Limits are explicit. Jev does not chat, does not write code from scratch, and does not explain its reasoning. It correctly tagged 9,857 as prime with 55 percent, but failed on a very large prime. It recognized a Shakespeare line, but still said yes when sweet became honey. It picked the wrong student in the flower-garden riddle. In chess it lost on material to Fable by plus 16 yet won on time; Jev moved in seconds while Fable spent 6 to 15 seconds per move and flagged.

From Inbox to Browser: Standout Demos

The most repeated demo is email triage. Riley Brown classified 500 emails by category, priority, spam score and reply need in seconds. Ryan Vogle ran 1,700 emails through 4.2 million input tokens for 18 cents, with 500k output tokens of JSON, and the dashboard showed about 10 percent spam. A terse message about being charged twice leaned billing 63 percent but not decisively; a positive love the service note collapsed to none with 100 percent confidence.

The same pattern routes support tickets: department, urgency and frustration as three parallel questions. A Hinglish refund request mentioning software and money was correctly teased apart between technical and billing and scored 31 percent urgent. As a model router, a one-prompt Claude agent used Jev to pick nano for a simple hello, a mid model for an app idea, and a balanced model at 95 percent for a heavier code request.

Browser automation is the showpiece. In a wiki race from DNA to My Little Pony, Jev finished 5 hops in 0.5 seconds versus 4 to 5 seconds for Haiku and Sonnet. A Zurich to London flight search finished in 7.1 seconds for 0.4 cents. Doom ran for an hour on millisecond decisions for about $7. A rebuilt Tesla full self-driving loop recreated in under an hour stopped at a stop sign and navigated, while Melee and StarCraft ran with real-time Jev decisions.

Everyday cases fill the gap. A smart-home intent to turn off all lights executed in 185 milliseconds. A fabricated fraud invoice first scored 85 percent true, then 94 percent after fraud signals were added. A trading loop deciding buy, hold or sell every second lives at jevtrader.vercel.app, but its creators warn it is not profitable. A YouTube clip finder ingested 1.1 million tokens and scored 17 moments in 3 seconds.

Jev in the Coding Loop: Linters, Smells and Armies of Browsers

In engineering, Jev shines as a cheap verification layer. One demo built a qualitative linter for comments; multiply value by 2 is accurate but useless. Scanning 150 comments took 9.3 seconds for 1 cent; scanning a whole codebase for 1,700 candidates was estimated at 57 cents. Another pass scanned 28 million tokens for code smells for $1.19.

Running thousands of browsers in parallel becomes economical. The BrowserUse team showed flight search plus adversarial suites that crawl a release trying to break it for pennies. One shop runs verification on every pull request and feeds only strong failures back to System 2. On the Claude Code side, picking the right skill among 182 dropped the wrong-skill rate from 17 percent to 7.3 percent with Jev, saving about 10k context tokens.

Caveats are discussed openly. Jev gives no reasoning trace, so the confidence score is the only audit signal. For high-stakes loops like trading, speed without cross-checked news is not enough. For code, cheap scanning surfaces candidates but cannot catch every smell; the value is triaging so expensive System 2 review focuses only on high-probability hits.

Access and Where to Put It in Real Life

Access is currently via the typesafe.ai waitlist and instantly through the Vercel AI Gateway. Early testers report approval in hours, with $5 of trial credit and a playground for the three primitives. Installation as a Claude Code or Codex skill is documented; a single command with an API key lets agents start using Jev immediately, and the cookbook hosts copy-paste examples.

The practical rule from the videos is simple: if you look at data and make a decision, put Jev there. Score inbound leads from a contact form, route a support ticket to the right team, run a pull request through a 100-question sieve, or score a log line from 0 to 3 and page the on-call only above 2.5. One example scored a graphic-design agency's form as a good-lead probability, letting the founder reply to Coca-Cola-level leads the same day.

Visualization: nodesdaily AI

Speed per Decision

  • Jev0.15 s
  • Flash LLM1.5 s
  • Frontier LLM10 s
Typical time for the same classification JSON; parallel scoring advantage.
FeatureJev (TypeSafe)Traditional LLM
ModeDecision engine, JSONChat, token stream
Latency100-200 ms4-10 s
Input price$0.042 / 1M$0.15-1.25 / 1M
Output pricefreepaid
Context64k1-1.5M
Hallucinationnear zerovariable

Key moments

  1. Intro: 200x faster decision engineChatGPT co-creator Diego Almeida reveals two years in stealth.
  2. RLCD vs RLHF: calibrated trainingOptimized for probability accuracy, not human preference.
  3. Three primitives: choice, score, boolEvery decision maps to type-safe JSON.
  4. Email triage: 500 and 1,700 emails in seconds18 cents classifies the whole inbox; spam share revealed.
  5. Browser race: 5 wiki hops in 0.5sJev beats Haiku and Sonnet by up to 10x.
  6. Games and autopilot: Doom, Minecraft, TeslaReal-time decisions stream for hours.
  7. Coding loop: linters and browser armiesComment scans and adversarial tests cost cents.
  8. Limits: large prime and logic missesSpeed alone does not replace reasoning; System 2 remains.

AI commentary

"The shortest way I can frame Jev is this: for three years we treated large models as chatbots, but the most expensive, slowest job was classification. Jev rips that job out and hands us a very cheap, very fast, very boring superpower."

AI assessment

The strength is picking the right lane. Instead of entering the text-generation race, Jev accelerates the dark work that repeats millions of times a day: classification, routing and verification. At around 150 milliseconds and cent-level cost, flows that were previously too expensive to automate become viable, especially support triage, email and pull-request checks. Even beating a small model in Vercel's test supports the calibration claim.

The limit is that calibration is not explainability. A score can be accurate while giving you no reason, which complicates auditing in regulated or high-stakes domains like finance and hiring. Failures on a large prime, a Shakespeare variant and a logic riddle show pattern matching rather than reasoning. Winning chess on time demonstrates how speed exploits tournament rules, not superior strategy.

At the industry level, Jev weakens the one-model-fits-all narrative. The future looks hybrid: System 1 scans fast, System 2 thinks deep and rewrites the rules. That aligns with the cheap-intelligence wave from open Chinese models: as intelligence cheapens, consumption rises, and the bottleneck shifts from tokens to compute and operations.

Pragmatically, it is worth piloting with managed expectations. Start with low-risk, high-volume loops — email, forms, logs and code comments — measure error against thresholds, and route only strong signals to an expensive model. Do not wire irreversible decisions like trading or hiring directly to automation; keep Jev as advisor and let a human or System 2 decide.

Sources

10 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

jev · typesafe · decision model · rlcd · system 1 · ai · automation

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…

What Is Jev? ChatGPT Co-Creator's 200x Faster Decision…