For years our relationship with AI was a prompt-in, text-out loop. Doruk Yalçınsoy marks a break from that: the model heard as ‘Cev’ in the video is actually Jev , TypeSafe AI’s first System One family member that opened in early access on 15 September 2026. The claim is that after instruction-following writers, a layer that reads a state and decides enters daily work. The demos are deliberately ordinary — swapping a garment on command, steering from streaming traffic data, sorting emails — all share one thing: not generating prose but filtering among bounded options with calibrated confidence .
Not a Writer, a Decision Engine
What sets Jev apart is architecture: three question types — Noul (yes/no probability 0–1), Choice (pick one from a list) and Score (graded 0–n) — evaluated in parallel and in isolation against the same state. Asking seven questions in one call barely moves latency because there is no text generation, only typed outputs with calibrated confidence . Under the hood TypeSafe describes ‘Reinforcement Learning for Calibrated Decisions’ and a new sampler. In practice that translates to 70–500 ms end-to-end, often around 100 ms, at $0.042 per million input tokens with free output , 32K context, text-only, gated by a waitlist while in early access.
The speed and cost claim lands with two numbers in the video. A marketing line of ‘4 seconds for 1,000+ emails’ measures as 6–7 seconds in Yalçınsoy’s own warm-up runs, about 300 ms per query . The same seven-question set on a large model like GPT 5.6 Luna takes roughly 5 minutes . TypeSafe’s own workflow evaluation cites 193.6× faster and 444.6× cheaper, with a clear footnote that it is vendor-run and reference answers were averaged from GPT-6 Astra and Fable 5.1. Even so, a back-of-the-envelope ‘500 million input tokens ≈ $21’ and free output make the case for splitting the decision layer onto a cheaper, faster engine easy to grasp.
From Inbox to Lead Scoring: Where It Fits
The clearest demo is the last 50 emails sorted for sponsorship relevance : 42 seconds total , 800 ms per verdict , with a final yes/no distribution in percent. Generalizing is straightforward: is this invoice or receipt , brand deal or phishing, newsletter or cold outreach; tone triage for YouTube and other social comments; post-meeting checks like ‘was a decision made, are action items defined, is follow-up needed’; sales-side checks such as whether a video clip can stand alone and its hook strength for virality estimation ; contract review — risky clauses, payment terms beyond 30 days, need for legal review. The video emphasizes lead triage — does the inbound have budget, does it match our audience, how urgent is it — and the advice to ask those three as separate questions rather than one vague ‘is this a good lead?’
Yalçınsoy’s daily stack is a console that bundles subscriptions — Cloud, ChatGPT, Gemini, Grok and Chinese models — behind one surface, with a harness behind it that serves all models together. The step described as ‘pick C and then choose cheapest tokens or highest quality’ is actually Jev as a router: given a task, it decides which model should handle it. In the example a blog post about ‘GPT1 Live’ is routed to GPT 5.6 Terra at 100% confidence and execution is delegated there. The routing signal comes from tables on arena.ee , described as human-voted model rankings; Yalçınsoy notes Turkish is missing from those tables and, from his own tests, keeps GPT as his pick for Turkish, shaping the prompt data accordingly. The same single-question Noul/Choice routing can be built for research, image generation or text-to-video categories.
Setup is kept intentionally quick. You craft a state plus questions at console.typesafe.ai/playground — Yalçınsoy even has another AI draft the questions because wording matters. Then you generate an API key , named ‘YouTube’ in the demo, copied with a blunt warning not to share it. Next you open Codex (or VS Code, Hermes, your console) and attach the key inside the project, ideally as an environment entry rather than pasted prose. Early access ships with $5 of credit and per-decision cost is framed as about $0.001 in the video; the test set covers three buckets — customer requests, content drafts, AI tasks — and is run as a tiny harness job. A practical note is decomposition: scoring a lead in components and combining in code yields more calibrated results than one monolithic question.
And the live test: a single sentence — ‘our checkout is broken, no customers can order, we are losing sales’ — asked as ‘Is this customer message urgent?’ comes back Noul 0.97 yes in about 0.8 seconds . Yalçınsoy’s lesson is the broader pattern: you do not need a giant model for every small verdict . The System One fast-intuitive layer (Kahneman’s System 1 metaphor) handles fast and consistent judgments, while large generative models are called only when the chosen path actually needs prose, planning or deeper reasoning. Limits are stated as well: Jev does not take image/audio/video input , accuracy can be lower on Japanese and similar inputs, a Noul near 0.5 means uncertainty, and it can stumble on counting or date math. Still, for repetitive, threshold-based work such as inbox triage, lead scoring, contract checks and content gating, the small decision engine plus large writer tandem is presented as today’s most pragmatic automation pattern.
| Topic | Summary |
|---|---|
| Model | ‘Cev’ in video is Jev — TypeSafe System One |
| Speed / Cost | 300 ms/query, 42s/50 emails, $0.042/MTok in |
| Use | Inbox, leads, contracts, comments + harness routing |
| Dimension | Jev (System One) | Large LLM |
|---|---|---|
| Input | Text-only, 32K | Text+image/audio |
| Output | Typed verdict + confidence | Free-form text |
| Speed | ~300 ms / query | Seconds–minutes |
| Cost | $0.042/MTok in, out free | $/MTok in+out |
| Best at | Triage, routing, scoring | Writing, reasoning, generation |
Key moments
- Intro — from prompting to deciding
- Jev’s three primitives: Noul, Choice, Score
Jev does not write, it returns typed verdicts
- Speed check — 6s vs 5min for 7 questions
- Inbox triage — 50 emails in 42 seconds
- Routing inside the harness
- Playground and API key setup
- Live urgency test — 97% in 0.8s
Checkout broken message marked urgent
AI commentary
"In my view Jev’s promise is not writing better but writing less. Instead of calling a giant text generator where a verdict is needed, placing a small engine that returns yes/no/category with a calibrated confidence score saves both time and budget across repetitive work like email, leads and contracts."
AI assessment
In my view the video’s strongest claim is also the one that needs the most careful reading: ‘don’t call a giant model for every small verdict.’ To steelman it: if you truly ask a large model about 1,000 emails one by one, moving to a typed, parallel engine like Jev at ~300 ms per query and ~$0.001 per decision is a decisive scale advantage, especially where no prose is needed — email, comments and lead triage.
What is missing is methodology and risk. A single sample of 50 emails, warm-up measurement and the empty Turkish column in the arena.ee tables weaken generalizability. The vendor-run 193.6×/444.6× figures are not independently verified; context width, threshold tuning and treating a Noul near 0.5 as ‘uncertain, needs human review’ are glossed over. The lack of image/audio/video input and lower accuracy on CJK inputs also narrows production scope.
I checked incentives and verifiability via web search: TypeSafe’s September 15 launch, 70–500 ms and $0.042/MTok pricing, 32K text-only and waitlist-gated access are consistent in docs, and independent summaries frame Jev as ‘decisions, not strings.’ The gap between marketing ‘4 seconds’ and measured ‘6–7 seconds’ is normal and stated transparently. Pricing, waitlist and language support can shift monthly, so console.typesafe.ai and docs.typesafe.ai need a live check at decision time.
My practical take: for small teams with repetitive, threshold-based work, Jev with $5 trial credit and teachable question design is the lowest-friction path today; for those about to hand production data and customer touchpoints to the cloud, do not skip human review points, thresholding around 0.5 and a two-week measurement. Yes for learning and triage, not yet as a stand-alone general assistant for everything.
Sources
7 links; 2 of them also cited by 3 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Doruk Yalçınsoy — Handing My Business Decisions to a New AI Model
- @docs.typesafe.ai https://docs.typesafe.ai/introduction
Also cited by: Building a Harness with Jev: LangChain and TypeSafe's System 1 Model · Jev: Not a Text Generator but a Decision Engine — TypeSafe's 70-Millisecond Move
- @docs.typesafe.ai https://docs.typesafe.ai/concepts/system-one
Also cited by: Nine Free AI Agent Skills Worth Installing Right Now · Jev: Not a Text Generator but a Decision Engine — TypeSafe's 70-Millisecond Move
- @oflight.co.jp https://www.oflight.co.jp/en/columns/typesafe-jev-system-one-model-2026
- @jevai.wiki https://jevai.wiki/
- @typesafe.ai https://typesafe.ai/
- @console.typesafe.ai https://console.typesafe.ai/playground
jev · typesafe ai · system one · decision engine · inbox triage