Almost every device you touch each day hides a central processor, from phones and laptops to game consoles and cars. The gap between generations is staggering: the 6502 chip inside machines like the Apple IIe used only 4,528 transistors and handled roughly 430,000 operations per second, while the Apple M1 packs about 16 billion transistors and reaches roughly 3 trillion operations per second. One ran plain early games and home software, the other renders rich 3D worlds and edits video with ease. According to wikipedia.org, the 6502 launched in September 1975 as an 8-bit microprocessor with 56 instructions.
The Shared Principle Behind Every Processor
The central claim is that both chips share the same technological DNA , despite about 45 years of progress between them. That DNA is not the transistor itself, nor the logic gate, but an architectural idea: every processor fetches an instruction, figures out what it means, and carries it out. Desktop CPUs, phone chips, graphics processors, and AI accelerators all inherit this pattern. The idea reframes progress as variation on one durable theme rather than a series of clean breaks. The low-cost MOS parts documented by commodore.ca disrupted the market with a famous 25-dollar price point and seeded that whole lineage.
Opening a modern laptop makes the abstraction physical. The tour passes the trackpad, battery cells, speakers, and cooling fan before settling on the motherboard, where SSD storage chips hold files durably and DRAM modules serve as short-term workspace. Under the heat pipe sit the memory and the processor package itself. That package stacks three elements: a metal lid that spreads heat, an interposer full of fine wiring and contact points, and the silicon die where the real computation happens. As announced on apple.com, the M1 arrived as a 5-nanometer chip with 16 billion transistors and unified memory shared across the processor.
Zooming onto the die reveals districts with distinct jobs: high-performance cores, energy-efficient cores, graphics cores, cache memory, and control fabric stitching them together. Dropping deeper, the view becomes a layered maze of wiring stacked over transistor beds at the bottom. One highlighted cluster shows six transistors wired into a single AND gate, with only about 650 transistors visible out of billions. The point lands visually: incomprehensible scale assembled from utterly plain switches.
To make that scale graspable, the explainer builds a library analogy and maps each fixture to hardware. The bookshelves are SSD storage, vast but slow to consult. The rolling cart is DRAM, a small working set wheeled near the desk. The desk itself is the CPU, and the single open book on it stands for cache memory , fast but tiny. The sheet of paper represents registers holding values in active use, the pencil reads and writes state, and the calculator is the arithmetic logic unit. Overseeing it all, a tireless control unit ferries books, reads pages, and operates the calculator.
How Programs Actually Run
The calculator deserves a closer look because its limits define what software can express. It works in binary, with only zeroes and ones, and performs plain operations: add, subtract, nudge a value up or down by one, and shift bits, which in binary doubles a value the way appending a zero multiplies by ten in decimal. It also applies bitwise logic such as AND, OR, and exclusive OR across each bit position. Most consequentially, it compares two numbers and raises flags for equal, less than, or greater than. Its display, which holds each result, earns its own title: the accumulator .
Running anything starts with loading: programs move from long-term shelves into the DRAM cart, then the active page lands on the desk like an open book. Each book mixes two page types, instruction pages numbered step by step and data pages pairing addresses with stored values. The walkthrough executes a Load that copies a value from a data address into a general-purpose register, an increment that routes the value through the calculator to add one, and a Store that writes the accumulator result back to the same address. Keeping place in this march falls to a special register called the program counter , which advances after every step.
Fixed sequences alone cannot express real software, so the next idea is the jump. A jump instruction writes a fresh number directly into the program counter, teleporting execution forward or backward instead of stepping onward. Conditional branches go further by testing the comparison flags first, which is how IF statements and loops emerge from hardware. The traced example runs a loop four times: load a counter, compare it against four, exit when greater or equal, otherwise run the body and jump back to the top. The readable text on screen is C++, the friendlier listing is assembly, and beneath both sits binary machine code.
Here the numbers surprise most newcomers. The little 6502 understood only 56 distinct instructions, while the modern M1 handles 354 under the ArmV8.4 set, and nearly every one is deliberately simple. Yet every game, editor, browser, and model response is ultimately some long sequence drawn from that short menu. Programs routinely chain millions of such steps, which is why a single wrong one can crash the whole edifice. Simplicity at the bottom plus enormous sequences on top is the whole trick of general computing.
The Cycle That Defines Every Processor
All of this condenses into the Fetch-Decode-Execute cycle , the heartbeat of the episode. During fetch, the controller uses the program counter to locate the next instruction, copies it into a holding spot called the current instruction register, and bumps the counter. During decode, an instruction decoder built from dense logic gates interprets the binary fields: which unit to use, which operation to apply, and which registers feed it. During execute, precisely timed electrical signals shuttle the register values into the calculator, let the gates settle, and latch the answer into the accumulator or back to memory. As codecademy.com explains in its instruction cycle guide, fetch, decode, and execute repeat for every single instruction the processor completes.
Speed comes from shrinking and overlapping that loop. The 6502 ticked at about 1 megahertz with 8-bit-wide registers, so each stage lingered roughly a microsecond. The M1 runs near 3.2 gigahertz on 64-bit paths, compressing stages toward a third of a nanosecond while moving far bigger numbers per step. Modern chips also pipeline: while one instruction executes, the next decodes and a third fetches, like an assembly line of overlapping cycles. A branch prediction unit guesses which way conditional jumps will go so the pipeline rarely stalls, rolling back on the occasional miss.
Beyond the Classic CPU
Not every chip plays by these rules. Mining ASICs and automotive FPGAs skip fetch and decode entirely, streaming data through fixed gates tuned for one repetitive job with superb efficiency and zero flexibility. Quantum hardware departs further still, trading bits for qubits and circuits of a different kind. The episode also names two often-omitted stages, Memory and Writeback, which shuttle pages between storage tiers and file results back, using helpers like the memory address register. These transfers are slow beside register arithmetic, which is why some textbooks fold them into the cycle and others treat them separately.
The finale climbs one layer up to instruction-set philosophy and chip layout. The Snake game demo compiles the same 145 lines of C++ into 676 ARM instructions on a RISC target versus about 560 denser x86 instructions on CISC, showing simple steady instructions against fewer but heavier ones. The single M1 core unpacks into split caches, an 8-wide pipeline, 32 general-purpose registers, and arithmetic split across eight smaller units with dedicated load-store paths. Graphics hardware scales the other way: thousands of simple CUDA-style cores all obey one fetched instruction on different data. The RISC and CISC tradeoffs debated on stackoverflow.com center on simpler fixed-rate instructions versus denser variable-length ones with complex decoders. For graphics hardware, cornell.edu describes the SIMT warp model where one instruction fans out across many threads with different data.
Key moments
- A 4528-transistor chip versus a 16-billion-transistor flagship
- MacBook teardown: board, storage, memory, and processor package
- Library analogy: shelves, cart, table, paper, and calculator
- Inside the ALU: binary math, shifts, logic, and flags
- Load, increment, store: a program takes its first steps
- Fetch, decode, execute: the loop behind every instruction
- Clocks, pipelining, and predicting the next branch
- RISC versus CISC and thousands of GPU cores in formation
AI commentary
"The library analogy is the standout choice here, making caches, registers, and the program counter feel concrete without dumbing them down. The constant shuttling between the tiny 6502 and the giant M1 keeps the scale gap vivid. As an overview it earns its length, though specialists will want deeper sources on pipelining and GPU design."
AI assessment
The strongest counter-view is that the Fetch-Decode-Execute loop gets more credit than it deserves. Memory specialists argue the real story of modern speed is the memory hierarchy, interconnects, and data movement, since arithmetic itself is cheap and waiting on data is expensive. Dataflow and domain-specific designers add that fixed pipelines waste energy on control overhead, which is why GPUs and AI accelerators push work into wide parallel lanes instead.
Several limits deserve a flag. The M1 block diagrams are acknowledged approximations because Apple keeps exact layouts proprietary, so core counts and cache sizes should be read as informed estimates rather than blueprints. Power delivery, thermal limits, cache coherence, and security side effects of speculation get little attention, and newer steps like Memory and Writeback are only sketched. The quantum section is a brief pointer, not a serious comparison.
The narrator is an educational animator with an obvious incentive to keep viewers watching through a long explainer, aside from a brief sponsor mention that stays clearly separated from the lesson. That incentive shows in the pacing: big visual payoffs arrive early, harder ideas like decoding and pipelining land later, and each section ends on a hook. None of that distorts the technical claims, which stay mainstream and checkable against standard references.
The practical payoff is a durable mental model for everyday performance. When a task feels slow, ask which shelf the data sits on: SSD storage, DRAM workspace, on-chip cache, or registers, since each step closer to the calculator is far faster. Developers gain intuition for why tight loops, predictable branches, and cache-friendly layouts win, and buyers can read chip announcements with cooler eyes by comparing memory systems and core layouts, not just gigahertz.
Sources
7 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Branch Education
- @wikipedia.org MOS Technology 6502
- @commodore.ca The Rise of MOS Technology and the 6502
- @apple.com Apple unleashes M1
- @codecademy.com Understanding The Instruction Cycle
- @stackoverflow.com How does the ARM architecture differ from x86
- @cornell.edu SIMT and Warp Cornell
Also cited by: 36 Trillion Operations: The Hidden Army Inside Graphics Cards
cpus · computer architecture · apple m1 · mos 6502 · gpus