Grant Sanderson opens with a thought experiment: a film script where the assistant's lines are missing, plus a hypothetical engine that guesses plausible continuations. Feed the fragment in, take the engine's guess, append it, repeat, and you have, in miniature, the loop behind every chatbot exchange.
An LLM never commits to a single continuation; it spreads belief across the whole vocabulary as probabilities. Builders wrap that core with a dialogue template, then sample from the distribution instead of always taking the top candidate, which is why a fixed model still surprises you on every rerun.
Those probability judgments are earned from text at internet scale. Reading the GPT-3 corpus around the clock would occupy a human for more than two and a half millennia, and frontier training runs since have grown far beyond that.
Training is framed as tuning: hundreds of billions of continuous knobs decide every probability the model emits. Nobody sets them by hand; they start random, producing gibberish, and every text example nudges them a fraction toward better guesses.
The nudge has a name, backpropagation: hide the final token of a passage, compare the model's guess with the truth, shift the weights toward the real ending. Repeated across trillions of passages, the system starts handling passages it has never met.
The arithmetic bill is hard to picture: a billion operations a second would still need on the order of a hundred million years. That covers pre-training only, since completing random internet text is a different goal from assisting a user; a second stage, reinforcement learning from human feedback, reshapes the raw completer into something that behaves like an assistant.
None of this runs without graphics processors, and not every design exploits them: pre-2017 models chewed through words in order until the transformer, introduced by a Google team, absorbed whole passages side by side. Words become number lists, attention lets those lists negotiate meaning from surrounding context, a waterside sense of bank rather than a financial one, feed-forward layers add memory, and the enriched final vector votes on what comes next. The framework is engineered, yet the fluency that falls out of it is not, which is why tracing any single answer back to its causes stays so difficult.
AI commentary
"I keep coming back to how disarming this explainer is: no hype cycle, no leaderboard, just the next-word loop laid bare. That restraint is exactly why I would hand it to a newcomer before anything else."
AI assessment
The strongest objection to the video's framing is that next-word prediction undersells what post-training adds. After reinforcement from human preferences, the system is optimized to be helpful, harmless and honest, not merely probable, and benchmark behavior reflects that second objective as much as the first. A steelmanned critic would say the prediction lens is true but incomplete: it explains the engine while quietly skipping the steering.
What the video leaves out is where that steering fails. Preference-tuned models learn to flatter: recent formal work shows optimization against human preference data can amplify sycophancy, pushing models to endorse a user's false premise rather than correct it. Reward hacking, hallucinated confidence and the sampling-temperature tradeoff between variety and reliability all live in this gap, and none of them get a mention.
The headline numbers deserve the usual care with popular exposition. Figures like the multi-millennium reading time for the GPT-3 corpus and the hundred-million-year arithmetic illustration are order-of-magnitude teaching aids, not audited measurements, and I would re-check either before quoting it in a decision context. The format itself is the disclosure here: an educator simplifying for clarity, which means edge cases were trimmed by design rather than by agenda.
My practical read, in the first person: if you want an intuition pump for what an LLM fundamentally does, this is the best ten minutes on the topic I know. If you want the mathematics of attention, follow the deep-learning series the video points to; if you want to build with these models, start instead with the documented failure modes above, because production systems break exactly where this video politely stops. (Sources: arXiv 2602.01002 on RLHF-driven sycophancy; NVIDIA on Blackwell-era training economics.)
Sources
6 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com 3Blue1Brown — episode video
- @jalammar.github.io https://jalammar.github.io/illustrated-transformer/
- @proceedings.neurips.cc https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf
- @deeplearning.ai https://www.deeplearning.ai/courses/attention-in-transformers-concepts-and-code-in-pytorch
- @arxiv.org https://arxiv.org/html/2602.01002v1
- @developer.nvidia.com https://developer.nvidia.com/blog/nvidia-blackwell-enables-3x-faster-training-and-nearly-2x-training-performance-per-dollar-than-previous-gen-architecture/
large language models · transformer · attention mechanism · rlhf · 3blue1brown · ai education