Is there a fundamental limit to how little data a text can be squeezed into? ASCII spends 8 bits per character; giving frequent letters short codes brings the average down toward 4 bits, and smarter methods that exploit patterns across long passages do better still. But where is the floor? The question reaches back to the 1940s, to the founding work of Claude Shannon that launched information theory, and the same mathematics now shows up in how language models are trained.
In today's language models, pre-training is described as next-token prediction with something called cross-entropy loss, a term rooted in information theory. One of the theory's conclusions is that prediction and compression are mathematically the same problem: predicting well means compressing well. That equivalence reframes the training objective for me: it is not really about token guessing, but about building the most efficient text compressor possible.
The phrase compression is intelligence is provocative, yet intelligence is such a stretchy, ill-defined word that the slogan alone cannot be judged precisely; the safer statement is that the mathematics of compression is strangely tied to artificial intelligence. Still, the short phrase provokes thought, so the series takes it seriously and spends three videos examining what the claim really amounts to. Part one's job is understanding the limits of compression and leading the viewer to reinvent the core idea behind Shannon's noiseless coding theorem.
The warmup example goes like this: a rover on a distant moon receives four movement commands, up, down, left, right, each a fixed-size step. The commands do not arrive equally often: half are up, a quarter down, an eighth left and an eighth right. For simplicity each command is assumed independent, drawn from this distribution regardless of context. The question: when the bit stream is slow and costly, what is the most efficient way to encode these commands as bits?
The straightforward student assigns two bits per command, which decodes easily but never exploits how common the up command is. The clever student proposes variable lengths: up becomes 0, down 10, left 110, right 111. The weighted average works out as half the time 1 bit, a quarter of the time 2 bits, an eighth each 3 bits, for 1.75 bits per command overall. That is a clear win over the flat 2-bit scheme.
Variable lengths raise a decoding puzzle: how does the robot know where the boundaries fall? In the proposed code no codeword starts another one, a property called prefix-free. The robot reads bits and records a command only once a complete codeword has formed. On the diagram of all binary strings, each choice consumes a share of code space: 0 takes half, 10 a quarter, 110 and 111 an eighth each. The shares match the command probabilities exactly, with nothing left over.
The theoretical student writes no actual code and instead ponders what properties an ideal code must have. The clever idea: random noise should be incompressible, so a perfect compressor must output a bitstream indistinguishable from random noise. A compressed message of n bits is one of 2 to the n possibilities, all equally likely under the noise appearance, so each source message carries probability 2 to the minus n. Encoding a message in fewer bits moves it down a layer of the diagram; a bit saved here costs two bits elsewhere, like pressing a bump in a carpet only to raise it worse somewhere else.
This reasoning makes the logarithmic expression unavoidable: a message using n bits in an ideal scheme has probability 2 to the minus n, so taking base-2 logs and negating equates bit count with negative log probability. Shannon named this quantity the information of an event: the column grows tall as probability squeezes toward zero and shrinks as probability nears certainty. When probabilities are not clean powers of 2 the value comes out fractional, and fractional bits make sense as a lower bound on whole-message code length rather than on any single symbol.
Language differs from the robot case in two ways: each new letter's probability depends heavily on everything before it, and the probabilities are never tidy powers of 2. A small on-device language model demo in the video produces a different distribution at every position, with the caveat that the model's numbers may not equal the true probabilities of language. A full sentence's probability is the product of successive conditional probabilities, the chain rule, and since logs turn products into sums, the whole message's information splits beautifully into per-symbol terms. Part three promises a concrete algorithm that compresses text to near that summed value.
One way to estimate the probabilities is scanning books and tallying what tends to follow short letter groups, but that collapses on long unseen contexts. Shannon instead played a guessing game with his wife Betty: she guessed a passage letter by letter while he wrote the correct letter on misses and a dash on hits. The abbreviated copy carried the same information, since her duplicate could regenerate the original by replaying the game. The 1950 paper scaled the experiment up, recording how many guesses each letter took and mapping that count to an implied probability.
The philosophical note is that Shannon was not doing pure data analysis; he was probing intelligent black boxes, treating interviewed brains as undescribable yet sophisticated language models. Today we build the black boxes instead of interrogating them. Entropy is defined at this point: the average information per symbol of a distribution, the weighted sum of each probability times its information value. The naming anecdote credits von Neumann, though the video itself notes the documentation is shaky.
The intuition runs like this: the more evenly a distribution spreads, the higher the total entropy, while one dominant event drags it low since the likely outcome carries little information; splitting probability across more symbols raises it. This is essentially the 1948 paper's noiseless coding theorem: no code beats this bound, and the bound can always be approached arbitrarily closely. For context-dependent processes like language the generalized version asks for the entropy rate, averaged over all possible messages. With at least a hundred letters of context, Shannon estimated English at about one bit per character, as if the language could squeeze into a single yes-or-no answer per letter.
Part two opens with a striking 2002 paper: Benedetto, Caglioti and Loreto showed in Physical Review Letters that gzip can recover structure between languages. The method appends a small snippet of document B to document A and compresses the result, then compares against compressing A alone. The difference measures how well the B snippet compresses under a dictionary tuned for A: small when the languages resemble each other, large when they differ. Language recognition, authorship attribution and the language family tree were all reportedly recovered with this co-compression distance.
The scene then returns to the robot: mission control changes the plan so up and down each drop to one eighth, left rises to a quarter and right to a half, while the decoders stay hard-coded to the old scheme. Weighting old code lengths by new frequencies gives one eighth of the time 1 bit, one eighth 2 bits, three quarters 3 bits, averaging 2.625 bits per command. That number is the cross-entropy of the old distribution relative to the new one, asking how a code tuned for one context performs in another. Formally, a Q-tuned code spends minus log2(q_i) bits per symbol, so under reality P the average is the sum of p_i times minus log2(q_i). In the bar diagram the widths carry P while the heights carry Q's information values.
Two-outcome toy distributions sharpen the intuition: with Q even and P skewed, every symbol carries one bit of information so the weights change nothing and cross-entropy stays at 1 bit. Flipped around, Q's entropy falls below one bit, yet the same code facing an even reality spends about 1.74 bits; order matters because P and Q play different roles in the formula. Fixing P and varying Q traces a curve minimized exactly where Q equals P, and that minimum value is the entropy of P. The green curve traced by those minima is, in the two-outcome case, the entropy curve of P itself.
The zipping trick rereads through this lens: compressing a B snippet with an A-tuned compressor is an empirical estimate of cross-entropy. Not literally: documents stand in for distributions, and gzip is far from Shannon-limit compression, replacing repeats with pointers via LZ77 before Huffman coding. Even so the distance measure does real natural-language work such as authorship attribution. The lesson generalizes: cross-entropy appears wherever the patterns of one setting must be measured against another.
The application is language model training: text is split into tokens, the model outputs a next-token distribution per context, and the loss averages the negative logs of the true tokens' probabilities. Machine learning uses the natural log, which differs from base 2 by a constant factor absorbed into the learning rate. Why the log at all? A pattern like my name is blank recurs thousands of times with different names; the model's Q distribution and the data's P frequencies define a total loss writable with a generic per-example function F of Q. Demanding the total be minimized only when the model matches the data forces, via Lagrange-multiplier optimization, the derivative of F into a constant-over-q shape that only logarithms possess. The hand is forced: the loss is not chosen, it is entailed.
The soft-target variant is distillation: the small model trains against the large model's full distribution at each point rather than the single true token, with cross-entropy written out explicitly. The analogy is chess: silently watching games versus a stronger player explaining how heavily each candidate move weighs at every turn. A footnote defines KL divergence as cross-entropy minus entropy, the per-symbol bits wasted by a mistuned code: zero when the distributions match, growing as they diverge, asymmetric. Part three promises the algorithm turning a predictor into a genuine compressor, equating pre-training with optimal-compression training.
AI commentary
"Watching these two videos back to back left me with one impression: the logarithm is not a loss function to memorize but the stop everyone must reach once the compression question is taken seriously. The series turns formulas from a memorization list into the unavoidable result of reasoning, and I structured this article in that same order."
AI assessment
Steelmanning the other side: compression is passive, while intelligence needs goals, action and sequential decisions; even a perfect predictive distribution never chooses what to do with its predictions. A zip file is not an agent, and this objection points at the decision-theoretic steering wheel inside Hutter's own definition of intelligence. If I had to defend the slogan, I would concede that compression supplies the world-model engine while the steering must be sought elsewhere.
There are limits the video leaves out: reported experiments never count the size of the compressor itself, so even when a 70-billion-parameter model squeezes foreign modalities impressively, adding the parameters makes the two-part description far larger than the raw file. The series also skips the memorization question: models memorize until capacity fills and only then generalize, while gzip sits far from the Shannon limit. These three points complete the frame for reading the video's numbers.
Whose claim and whose interest applies here too: the talent page and the KL puzzle at the episode's end connect organically to the content yet remain a commercial showcase. Figures like the Chinchilla rates and the minus-0.95 correlation arrive via secondary retellings, so at decision time the DeepMind paper's setup and the correlation study's span of 31 models over 12 benchmarks deserve independent checks. In this article I used such numbers as signposts, not foundations.
My takeaway is this: for someone who reads loss curves, decides on distillation, or wants information theory by intuition, this series is a first-class stop; for anyone assuming lower loss automatically means a smarter model, it is a trap. Loss is a proxy metric, not a capability certificate. I take the compression-intelligence link seriously, but I always write the equals sign between them in pencil.
Sources
12 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com Episode video (Part 2)
- @youtube.com Episode video (Part 1)
- @3blue1brown.com https://www.3blue1brown.com/lessons/cross-entropy/
- @3blue1brown.substack.com https://3blue1brown.substack.com/p/reinventing-entropy
- @3blue1brown.substack.com https://3blue1brown.substack.com/p/but-what-is-cross-entropy
- @pubmed.ncbi.nlm.nih.gov https://pubmed.ncbi.nlm.nih.gov/11801178/
- @arxiv.org https://arxiv.org/pdf/cond-mat/0108530
- @math.mit.edu https://math.mit.edu/~shor/18.310/noiseless-coding.pdf
- @sebastianraschka.com https://sebastianraschka.com/faq/docs/next-token-prediction.html
- @arxiv.org https://arxiv.org/abs/2309.10668
- @arxiv.org https://arxiv.org/abs/2404.09937
- @shanakacdesoysa.substack.com https://shanakacdesoysa.substack.com/p/compression-isnt-intelligence-so
entropy · cross-entropy · compression · information theory · shannon · 3blue1brown · language models