Back to feed

Decisions as Retrieval: CLM-8B, the 13x Faster Architecture

Built jointly by Stanford and Nvidia Research, CLM-8B is an open 8-billion-parameter model that moves decision-making from text generation to retrieval. It measured 44 milliseconds against 570 at 1024 options with similar accuracy, though hits fade clearly as the option count reaches the thousands.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — eSuMmMMMrm0
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Most decisions inside real software do not need a model to think for thirty seconds. Is this ticket urgent, which tool should run next, which link should the agent click: those answers are needed in milliseconds, not seconds or minutes. That is why the System One wave keeps growing; fast, non-reasoning models that return typed decisions instead of essays are suddenly hot again. The opening thesis here is exactly this: the heart of agent loops is no longer generation, it is selection speed.

This week the difference is architectural. A joint team from Stanford and Nvidia Research released CLM-8B, an open 8-billion-parameter model they call a contrastive language model . The sharpest analogy is CLIP for decisions: the joint embedding idea described in the Arxiv CLIP paper is applied here to situations and actions. A state encoder reads what is going on, an action encoder reads what could be done. In the Mario example the game screen is the state and left, jump, or run are the actions; every option is embedded, scored by closeness to the state, converted to probabilities with a softmax , and jump wins at eighty percent.

CLIP for Decisions: State and Action in One Space

The clever part of the design is that it speaks the existing habit fluently: the model uses the same interface as Jev. According to the Jevapi docs there are three primitives; a yes-no question, a multiple-choice question, and a score. According to the System-One docs all three are the same thing underneath: a state plus a closed candidate list. That closed nature cuts both ways. The model cannot invent an answer outside the list, echoing the no-hallucination promise, but the Jevals methodology reminds us what a System One model is by definition: no out-of-list generation, yet in-list mistakes always possible, and the model can pick the wrong option confidently.

How is a space built where the right action lands closest? The answer is contrastive training . Training starts from millions of examples, each a pair of a situation and an action somebody actually took; the model learns to pull each situation toward its own action and away from the rest. The neat trick is that nobody writes down wrong answers: in a batch of a thousand, if jump is right for one Mario screen, the other 999 actions automatically count as wrong for it. Positive and negative balance comes for free, and the difference between sounding right and being right gets carved into the geometry of the space.

The secret of the speed hides in that split. In a Jev-style model the state and all options are read together on every call; doubling the options doubles the reading load, plus the answer must be decoded token by token. CLM separates the two and exploits the typical agent loop: what changes across steps is the state, not the action list. Actions are embedded once and cached ; each new step costs one state embedding plus a nearly free dot product. On the team chart the CLM line barely moves as options approach the thousands: 44 milliseconds against 570 at 1024 options, roughly a 13x gap. According to VentureBeat the team reports up to 9x in open tests; whichever number is taken, the chart points the same way.

Three-Stage Training: World First, Fine Distinctions Second

Training follows a three-stage recipe, and the order is the whole trick. Stage one feeds 60 million question-answer pairs from Nvidia Nemotron data for general world knowledge, with questions as states and answers as actions. Stage two adds 30 million hard negatives written by a generative model, answers that sound right but are wrong, like the Venus trap for the planet closest to the sun. On a hard-negative test, pretraining alone scores 52 percent while the second stage lifts it to 69; starting with hard negatives from day one peaks at 62 and slides into memorization. The Notion official page frames this order as a pretraining plus post-training relationship: learn the world first, then the fine distinctions. Stage three converts the model into an agent with about one million real trajectory steps, replaying 40 percent of the original data; without that replay the hard-negative score falls from 69 to 56.

The most surprising part is what never moves: the 8-billion-parameter backbone stays frozen as a Qwen3-8B reader, receiving not a single gradient. Only small heads of about 20 million parameters on each side are trained. Because the backbone is fixed, the dataset is embedded once and full pretraining finishes in about an hour on a single RTX 4090. The Qwen3 repository on GitHub confirms the open-weight, 36-layer structure, and the model card on HuggingFace records 8.2 billion parameters under Apache 2.0. Loss follows a power law in compute, data, head size, and encoder size, with the encoder delivering the largest gains. Hence the team points to the bigger multimodal release expected next month.

After theory the focus moves to a real desk: a Qwen3-8B encoder served with vLLM exactly as the Qwen docs describe, topped with the released 75 MB CLM server, for roughly 25 GB of total memory. According to Nvidia the DGX Spark holds inference up to 200-billion-parameter models on the desk with 128 GB of unified memory, so the rig is ambitious yet reachable. The state is a support ticket from a customer charged twice and unable to reach anyone; is it urgent, which team, how angry: answers arrive as probabilities, 84 percent urgent and 99 percent billing. Once the ticket is seen, repeat questions return in under 2 milliseconds because everything is cached. In the Romeo and Juliet test raw embeddings rank Marlowe first while CLM pushes Shakespeare to 98 percent; twenty million parameters bend the space from similar to right.

Speed at a Thousand Options, Limits on Accuracy

The architecture shows off at extreme scale: with a list of 1080 tools from Stripe to Trello, the first call takes seconds for embedding, then every later request settles near 80 milliseconds. Picking among a thousand costs nearly the same as picking among one. In wiki racing, a page with 934 links is scored against the target in a single pass and reached in four clicks. But the same experiments document the ceiling: correct picks fall from 86 percent with 8 tools to 17 percent with 1080, as lookalike refund tools get confused. The cause is the blind spot of the design: the state is encoded alone, so options are never laid side by side and compared; each candidate is judged in isolation.

Hence the two-layer advice to builders: let CLM shrink a thousand options to a ten-item shortlist in milliseconds, with a joint-reading model issuing the verdict. Where the action set is small and stable, a game loop or a few routes, the model does fine on its own. According to the tool-use guide from MLflow, piling more than 8 tools onto one agent is a design smell and splitting into sub-agents by function is the fix; the accuracy curve of CLM independently confirms that rule. The closing big picture is nostalgic: the fast no-reasoning idea that Jev reheated meets the contrastive algorithms of 2017 and 2018 in the same pot. Decisions stay with retrieval, generation with slow thinking.

Visualization: nodesdaily AI

Key moments

  1. System One wave: millisecond decisions
  2. The CLM idea: CLIP for decisions
  3. Speed chart: 44 ms versus 570 ms
  4. Three-stage training and hard negatives
  5. Ticket demo and Shakespeare test
  6. Limits at 1080 tools and wiki racing

AI commentary

"The speaker does not leave theory on the table; he runs the model on his own hardware and pushes it to the limit with a thousand-tool list and a wiki race. His honesty about the measured accuracy drop makes the 13x speed claim credible. To me the real value is the recipe, not the architecture: let the fast model shortlist, let the careful model decide."

AI assessment

The strongest objection is architectural, not empirical: because the state is encoded alone, the model never compares options side by side. Each candidate is judged on its own, so lookalike choices blur together and accuracy falls as the list grows. Retrieval is not reasoning, and similar is not the same as right; the slide from 86 to 17 percent is that sentence expressed in numbers.

What is missing matters too. This release is text-only, with the larger multimodal version still pending, so any claim about general decision-making waits for that test. The headline figures come from the team charts plus one set of demos, and independent replication has yet to arrive. There is also a quiet circularity worth noting: hard negatives generated by a large model teach the boundary between plausible and correct, but they also inherit that generator's blind spots.

Consider the presenting position with a cool head. The demos run on Nvidia hardware with an enthusiastic, builder-friendly tone, and excitement is part of the format. Against that, two facts discipline the story: the code and the server are openly released, and according to VentureBeat the project lead Jacky Kwok frames the contrastive objective as an advantage for domain-specific decisions rather than a universal victory. Open artifacts plus a bounded claim deserve more trust than either alone.

For the reader the practical lesson is concrete. If the action set is small and stable, a game loop or a handful of routes, this model class works well today. At larger scale deploy it as an opening filter that narrows a thousand candidates to ten within milliseconds, reserving the verdict for a model that compares the options jointly. And keep the training moral from the replay experiment: when adapting any model, keep mixing in the old data or the earlier gains quietly evaporate.

Sources

11 links; 2 of them also cited by 2 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

clm-8b · system one · contrastive learning · agent tools · qwen3

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…