Imagine a single box performing thirty six trillion operations every second and try to feel the scale. In 1996 Mario 64 needed only one hundred million operations per second to run smoothly, which looks tiny today. By 2011 Minecraft demanded one hundred billion operations, lifting the bar a thousand times higher. Modern titles such as Cyberpunk 2077 push the same scale another thousand times further. The comparison used to explain it is striking: we would need 4400 Earths full of people each doing one multiplication per second.
Inside this card more than 10752 CUDA cores work together, while a desktop processor nearby offers only 24 cores, and the contrast feels extreme. The explanation compares a GPU to a massive cargo ship and a CPU to a jumbo jet, trading huge capacity at slower pace against small load at high speed. Flexibility favors the processor, because graphics units cannot run an operating system or handle network connections directly. These core and bandwidth numbers broadly agree with the specification sheets published by Nvidia for this generation.
At the center sits the GA102 die with 28.3 billion transistors, most of its area devoted to compute hierarchies. The die splits into 7 graphics processing clusters, each holding 12 streaming multiprocessors that branch further downward. Every warp contains 32 CUDA cores plus one tensor core , reaching 10752 CUDA, 336 tensor and 84 ray tracing cores across the chip. This hierarchy and core breakdown is also described in similar form within the GA102 listings maintained by TechPowerUp.
Cargo Ship or Jumbo Jet
Curiously the 3090 Ti, 3090, 3080 Ti and 3080 share the same GA102 design, with luck during manufacturing deciding the label. Dust or patterning faults can damage a small region, so engineers isolate the faulty block instead of discarding the whole die. Flawless dies become 3090 Ti cards, while 3090 keeps 10496 cores, 3080 Ti keeps 10240 and 3080 keeps 8704 cores. The boards also differ in clock speed, memory size and memory generation, which further separates price and performance tiers.
Zoom into one CUDA core and you find a tiny calculator built from roughly 410 thousand transistors. The busiest section performs A times B plus C, called fused multiply add , which dominates graphics workloads. About half the cores favor 32 bit floating point values, while others handle integers or mixed formats, with extra logic for signs and bit operations. Division and square roots move to special function units , only four per multiprocessor. With 10496 cores at 1.7 GHz the math reaches 35.6 trillion operations per second, a remarkably consistent figure.
Stepping back, the board combines circuit, power stages, cooling and memory into one system. The regulator drops 12 volts to about 1.1 volts and feeds hundreds of watts to the processor. A cooler with heat pipes spreads warmth to fins where fans exhaust it. Game scenes move from SSD into 24 GB of GDDR6X memory , with a 384 bit bus at 1.15 TB per second versus 64 GB for desktop memory. This capacity and bandwidth approach is likewise presented in the same direction by Micron in its GDDR6X product notes.
Defective Chips Are Not Thrown Away
Moving that much data cannot rely on simple zeros and ones alone, so engineers play with voltage levels. Older GDDR6X uses four level PAM-4 signaling for two bits per symbol, while newer GDDR7 shifts toward three level PAM-3 coding. Small groups map into ternary symbols, longer blocks compress further, and useful data per wire rises clearly. For AI accelerators, stacked HBM3E cubes with TSV links reach 192 GB nearby. These PAM and ternary details are also summarized plainly in the pulse amplitude modulation entries on Wikipedia.
Memory Hunger and Signaling Tricks
The magic comes from single instruction multiple data , applied to problems that split with almost no effort. A cowboy hat uses about 28 thousand triangles and 14 thousand vertices defined in its own model space. Placing the scene adds each object position to its vertices, converting model coordinates into shared world coordinates. The scene repeats this across 5629 objects, 8.3 million vertices and 25 million additions. Since no calculation waits for another, thousands of cores share the work and assemble the world quickly.
One Instruction Across Millions of Data
Hardware organizes this through threads, warps, blocks and grids, each level managing a different scale. Thirty two threads form a warp sharing one instruction stream, warps gather into blocks, and blocks gather into grids. A Gigathread scheduler assigns blocks to units and balances traffic intelligently. Older designs moved warps in lockstep, while newer SIMT gives every thread its own counter and softens divergence with 128 KB shared L1. This lockstep history and loom inspired naming is explained in similar terms by parallel computing notes archived at Cornell.
The same parallelism helps beyond games, with two striking examples closing the story. Bitcoin mining repeats SHA-256 hashing while varying a random nonce, and this card produces about 95 million attempts per second. Yet dedicated ASIC clusters reach 250 trillion hashes per second, making graphics cards look like spoons beside excavators. Tensor cores instead multiply two matrices and add a third, repeating that template trillions of times for AI. Ray tracing cores get only a brief mention here and point viewers toward a dedicated episode.
Key moments
AI commentary
"The narrative moves cleanly from scale to silicon, balancing numbers with memorable comparisons. Binning and signaling sections teach the most, while the sponsor passage stays brief and tolerable. It works well as an entry point for curious hardware readers."
AI assessment
Credit the strongest counterargument first: a graphics processor is not a general brain and falls behind a desktop processor on any flexible task. Energy matters too, because hundreds of watts plus heavy cooling are required to sustain this pace. The Bitcoin example proves the limit, since dedicated ASIC clusters do the same work thousands of times more efficiently and pushed graphics cards out of mining.
The gaps are visible as well: ray tracing cores are only flagged in passing, with no clear operations per watt figure for efficiency. Memory sections lean heavily on Micron, and praise for GDDR and HBM drifts toward promotion. That slant does not falsify the facts, but readers should keep the sponsor balance in mind.
The takeaway for readers remains useful: parallelism, memory hunger and die sorting together explain modern speed. Huge memory paths serve games, matrix engines serve AI, and trial abundance serves mining, all from one principle. After this piece, the next loading screen reads more clearly as data moving between SSD and VRAM.
Sources
6 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Branch Education
- @nvidia.com Nvidia RTX 3090 family specifications
- @techpowerup.com TechPowerUp RTX 3090 database entry
- @micron.com Micron HBM3E product page
- @wikipedia.org Wikipedia GDDR7 memory article
- @cornell.edu Cornell GPU architecture notes on SIMT and warps
Also cited by: How Every CPU Thinks: From the 6502 to the Apple M1
graphics card · gpu architecture · cuda cores · gddr6x · artificial intelligence · bitcoin mining