Back to feed

MIT Calculated the Scaling Wall: Why Wider LLMs Stop Getting Smarter and Why Money Can't Fix It

A new NeurIPS study shows why large-model error falls as 1/width and why that rate is locked by Zipf's law. Doubling width halves interference and error, yet multiplies training cost by six, and even infinite money cannot lift the ceiling.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — 6xQ8LQfkBg4
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

The video opens with a familiar story: GPT-1 looked like a tiny dot, a year later GPT-2 was visibly larger, and then GPT-3, the engine behind the first ChatGPT, dwarfed them both. Performance climbed with size and the industry learned one slogan: bigger is better, much bigger is much better. Models train with self-supervised learning, a method that makes data labeling cheap and lets teams ingest all of Wikipedia, hundreds of thousands of books and a filtered copy of much of the public internet at once. From that enormous reading, without anyone teaching a single grammar rule, they build a huge mathematical map that encodes how every word relates to every other word.

Why width is scarce: 50,000 concepts, a few thousand dimensions

In an idealized model each of roughly 50,000 words and concepts, a conservative floor, would own a fully independent direction in the internal geometry. In practice the width that researchers call dimension stays at a few thousand, a little over 12,000 even at GPT-3 scale. We once assumed that when space ran out the model simply dropped low-value concepts. Anthropic's 2022 toy model showed the opposite: the model squeezes concepts into overlapping directions instead. This is superposition. In practice, banana, elephant, bunk bed, balloon and the semicolon can end up pointing almost the same way. The trick works only because not every concept is needed at the exact same instant in a sentence, much like not every item in a crowded closet is worn at once.

Overlap is not free. When the model tries to fire the elephant concept, neighbors that share the same direction fire faintly as well and meanings bleed into each other. Researchers call this bleed interference. Interference explains why a model sometimes picks a slightly off word and sometimes a wildly wrong one. The chain is three steps: 1) concepts are squeezed into limited width, 2) angles between vectors shrink, 3) co-activation creates crosstalk. Each step adds noise to the word map.

One formula: double width, halve error

The MIT result highlighted in the video is crisp: double the width and interference halves, and because interference drives error, error halves as well. A wider mathematical space spreads the same crowded concept set more thinly, which is a large part of why bigger models look better. The industry has pulled that lever for six years. Yet spreading has an end. Once width is large enough to give every concept its own direction with zero overlap, there is no superposition left and therefore no interference left to reduce. Beyond that point making the model wider buys nothing; that is the ceiling.

Even before the ceiling, three brakes appear and the first is money. Halving interference needs doubling width, but each doubling multiplies training price by six in the conservative estimate. If a frontier model costs about $300 million today, one doubling is about 1.8 billion, the next about 11 billion, then about 65 billion, then about 390 billion dollars after four doublings, a scale that is absurd even for OpenAI or Google. Imagine energy became free and money stopped mattering; the fall rate still would not change. The exponent is locked by human language itself. Zipf's law says a tiny set of words appears extremely often, such as that and is, while a huge set such as photosynthesis or kaleidoscope appears very rarely. That skewed spread locks the 1/width exponent in place and blocks faster improvement.

The error that never reaches zero and the absolute ceiling

The second brake is the nature of language: it is not an equation where 2 plus 2 equals 4. In everyday speech after I really like the next word could be cats, dogs or pizza with no single correct answer in the training data. That intrinsic uncertainty prevents total error from extrapolating to zero. Researchers write loss as a width-dependent term plus a constant and find that real model loss sits close to a straight line in 1/width. The Chinchilla picture fits: with parameter count scaling as width to the 2.52 power, the exponent versus width comes out at 0.88 to 0.91, close to 1. No matter how large the model grows, once the interference part shrinks, the language uncertainty floor remains.

The threshold that makes the ceiling concrete is when width catches the number of truly independent items in the vocabulary. Roughly 50,000 independent directions are needed, but we have about 12,000, so superposition is heavy. In toy experiments the paper distinguishes regimes: in weak superposition loss follows a power law only when feature frequencies themselves follow a power law; in strong superposition loss comes from geometric overlap and follows 1/width robustly across many frequency families. Measured on open models, mean squared overlap of normalized rows of the final projection matrix scales as 1/width, and those models sit in the strong regime. So the day width matches vocabulary size, superposition ends and extra width is wasted. Until then every doubling gives half the error, but the price for that half compounds.

Beyond text: models that see the world

If the scaling wall holds, what comes next? The video points to models that learn from video and senses rather than text alone, often called world models. The intuition mirrors predictive coding in the brain: predicting the next frame forces a model to grasp gravity, fluid mixing, solid collisions and similar physics. OpenAI's Sora implements this by turning video into spacetime patches; where large language systems have text tokens, Sora has visual patches that act as scalable tokens. Yet the Physics-IQ benchmark of 396 controlled videos shows that current generators, Sora, Runway, Pika, Lumiere and VideoPoet, still lack deep physical competence and that visual realism does not imply physics understanding; the best score is 29.5 against a real-world variability baseline of 100. Synthetic data hints that physical laws can be learned from observation, but passive watching without interaction misses causality. That is why new proposals try to encourage superposition deliberately, for example nGPT which constrains states and weight rows to the unit sphere, to make smaller models more efficient, or explore interactive training.

Visualization: nodesdaily AI

AI commentary

"My take is simple: the scaling mantra we memorized for years just hit mathematics. Wider is not always better; geometry, language statistics and the budget brake at once. The next leap will not come from more text width, but from models that actually see the world."

AI assessment

Steelmanning the other side: scaling advocates are not wrong. Chinchilla-style analysis shows an optimal balance among parameters, data and compute, and that data scale also lowers error alongside width. The ceiling cannot be debated on width alone; smarter data curation, better optimizers and architectural tweaks can deliver more for the same budget. In that frame, designs that deliberately encourage superposition, such as nGPT which constrains states and weight rows to the unit sphere, promise similar performance at smaller width for the same loss and try to route around the wall.

Limits matter. The bridge from toy model to real language-model head is insightful but narrow. Experiments focus largely on rows of the final projection matrix; the skills, grammar and reasoning learned inside transformer layers are not fully captured. The constant uncertainty term tied to evaluation set and model family forms a floor that never hits zero, and separating representation error from modeling error remains hard. The finding that current open models sit in strong superposition does not rule out future architectures that shift the regime.

On verifiability the picture is solid but confined. The paper is Superposition Yields Robust Neural Scaling at NeurIPS with open code, and the 1/width scaling of mean squared overlaps is reproducible. Still, large-scale claims need independent labs to repeat measurements across different vocabularies and languages at matched widths. The video narrative carries an MIT label and builds its cost projection from a $300 million baseline to about $390 billion after four doublings, relying on the conservative sixfold per doubling assumption and ignoring energy, hardware efficiency or algorithmic jumps.

The practical takeaway is to stop expecting miracles from scale and to recalibrate. For builders the message is to design more efficient rather than merely wider models; an architecture and training regime that manages superposition well at small width can win on cost. For enterprise buyers the table is a reminder that each sixfold budget increase for a frontier model halves error at best and does not erase the irreducible uncertainty of language. For end users the near future lies not in thicker text models but in world models that understand video and interaction; in the short run the best bet is the right narrow model plus a strong data pipeline.

Sources

7 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

llm scaling wall · superposition · zipf law · interference · world models

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…