Back to feed

AI's power bill: moving data now costs more than computing it

Data centers will double their power draw by 2030, and over 80 percent of energy in the largest AI clusters goes to moving data rather than computing. Huawei's UnifiedBus architecture promises to cut that waste with one protocol and shared memory.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — qJvE_RbxPjo
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Data centers are about to swallow more electricity than all of Japan, reaching roughly 945 TWh a year by 2030, with AI as the main reason. The striking detail is where that power goes: in the largest AI clusters more than 80 percent of energy is spent moving data between chips rather than computing on it. This 945 TWh projection comes from the IEA's Energy and AI report base case, which expects data-center demand to grow several times faster than all other sectors combined. Everything that follows is really the story of this single imbalance between data movement and calculation.

Frontier laboratories now build models with up to 5 trillion parameters , and most experts expect 10 to 40 trillion within a couple of years. No single chip can hold such a model, partly because of physics: the machines that print chips expose a rectangle of about 26 by 33 millimeters at a time, which caps how large one die can be. That is why powerful processors are assembled from multiple dies, and why memory must scale together with compute. Without enough RAM a giant model cannot even be loaded with context for all its users, so thousands of processors must work from a shared memory pool as one machine.

Why one chip is never enough: the physics and memory walls

Computer networking was solved long ago, but for a different job. Most of the world's traffic runs on TCP/IP , a design from the seventies whose routers forward packets and quietly drop whatever they cannot handle, leaving reliability work to the software at both ends. That division of labor scaled the internet beautifully, because a streaming video survives on buffers and nobody notices a late packet. AI training enjoys no such luxury: thousands of chips advance through the same step together, then each shares what it learned and waits for everyone else, so every round trip of tens of microseconds repeats millions of times and becomes the tax on the entire system.

Today's answer is a two-tier network: very fast links inside the rack and much slower ones across the data center. Nvidia's NVLink reaches up to 1.8 terabytes per second within a rack, yet once data leaves the rack the speed falls toward roughly 200 gigabytes per second, a drop Huawei's paper calls a cliff. According to Nvidia's developer blog, sixth-generation NVLink offers 3.6 TB/s per graphics processor and beats conventional Ethernet by a wide margin at rack scale with in-network computing. Smoothing that cliff into a gentle slope is exactly the premise of the new architecture.

One language for the whole data center

The first requirement is a single protocol from inside the chip package all the way across the data center, so no border crossing ever repacks the parcel. Every conversion between today's half-dozen link standards adds latency, hardware, and wasted energy, precisely where most movement power is lost. According to Huawei's official announcement, the wireless LinkBlade cabinet design removes about 196 kilometers of copper cabling inside a 4,096-NPU SuperPod. A small correction is in order: the presenter pronounces the architecture as Pirium in the video, but Huawei's official name in its announcement is UnifiedBus.

The second requirement makes distant memory feel local by giving every byte in the system a single address. Today a processor needing a number from another rack sends a request message, both network cards wrap and unwrap packets, and even streamlined paths remain message-and-wait systems. With the new approach the processor simply issues a load instruction , the same one it uses for its own memory, and pulls the data itself. Huawei puts the round trip at around 100 nanoseconds against tens of microseconds before, roughly 500 times faster, turning a bureaucratic exchange of requests into direct access.

The third requirement removes the master bottleneck that clogs large systems. Nearly every computer you have owned has a central processor issuing orders while accelerators and drives wait to be told, which works fine at small scale but jams as the army grows. The concrete AI example is the KV cache , the store holding everything a model has already read in a conversation. On the unified fabric a neural processor reads that cache straight off networked storage with no processor in the middle, so every millisecond saved fetching a long conversation's history is time before the answer starts appearing.

Copper ends, light begins, failures persist

The fourth requirement blends copper and optical cabling under one language, because copper has a hard limit at AI speeds. Each doubling of wire speed roughly halves the distance copper can carry the signal, since faster switching turns more energy into heat, leaving about one meter of useful copper reach. Pluggable optical modules keep a meter of copper in front of them, so Huawei chose near-package optics with only centimeters of copper beside the chip as the pragmatic middle ground. According to HPCwire's scale-up interconnect analysis, copper stays efficient over short reaches while rack-to-rack scale-out links make the shift to optics mandatory.

Light brings its own trouble: modulators and detectors add noise, with a linear optical link erring about once per 100,000 bits against Ethernet's expectation of fewer than one in 100 million. The usual fix of forward error correction adds tens of nanoseconds per hop, so Huawei skips it and lets each link check its own traffic and resend damaged data hop by hop. Optical links also flicker for milliseconds at a time, long enough to crash a two-month training run, and dual fibers on each cord stretch the average failure interval from about a day to about a month. According to SemiEngineering's CPO coverage, co-packaging photonic engines with the processor shrinks signal paths from centimeters to millimeters and cuts energy per bit.

The most striking concept of the trip is MFU , model FLOPS utilization: the ratio of a system's theoretical compute to what it actually delivers. At 30 percent efficiency a 4,000-chip installation performs like barely 1,200 chips, which is why faster silicon alone misses the point when the communication layer stays congested. According to ZeroEntropy's MFU concept page, the ratio divides achieved throughput by theoretical peak, and 40 to 60 percent counts as good for 2026 pretraining. Faster chips raise the top speed, but only faster and more direct communication clears the traffic, and an idle chip still burns power while waiting.

The showcase machine is the Atlas 960E SuperPod with 4,096 neural processors, using 5,500 optical engines instead of 48,000 plug-in modules and cutting one machine's draw by more than 550 kilowatts. According to Huawei's Atlas announcement, the first NPO-based SuperPod packs 4,096 NPUs delivering 8 EFLOPS in FP8, with 5,500 Hi-ONE units replacing 48,000 optical modules. The rollout calendar is staggered: version 1.0 has shipped since March 2025, version 2.0 arrives on the Atlas 950 late this year, the 960E has no published date yet, and its Ascend 960 chips target 2027. According to Huawei's IEEE ISCAS 2026 presentation, the Tau scaling law proposes scaling in time rather than geometry, with Logic Folding introduced as its first product. To speed adoption Huawei opened the specification under license, reporting 30,000 enterprise downloads across 26 countries plus an Ethernet-compatible variant.

Visualization: nodesdaily AI

Key moments

  1. The 945 TWh shock
  2. From 5T toward 40T parameters
  3. TCP/IP and classic networks
  4. NVLink and the cliff
  5. UnifiedBus and LinkBlade
  6. Distant memory in 100 ns
  7. Copper, NPO, and optics
  8. Atlas 960E and open license

AI commentary

"The numbers are dramatic, but the messenger matters: this is a vendor-hosted story, and the savings still need independent measurement. Read it as a credible direction, not a closed case."

AI assessment

The strongest counterargument is that none of the headline savings have been verified outside Huawei so far. The 550 kilowatt figure, the doubling of fault-free run time, and the 500-fold latency improvement all come from company materials and a hosted lab visit. Competing ecosystems such as NVLink are entrenched in real deployments with published third-party benchmarks, while UnifiedBus still has to prove itself on customer workloads rather than showcase machines. Until independent operators run the same training jobs on both fabrics and publish the power meters, the efficiency gap remains a claim, not a fact.

There are also gaps the video barely touches. Opening a protocol specification is not the same as building an ecosystem: switch vendors, optical module makers, server builders, and operating system maintainers must all invest before a single standard pays off. The video admits this buy-in problem openly, yet glosses over how long such coordination historically takes and what happens to early adopters if momentum stalls. Reliability figures based on dual fibers sound reassuring, but month-scale MTBF promises need years of fleet data, and the failure behavior of near-package optics at full data-center heat is still unproven at volume.

The presenter's interest is visible throughout: he flew to Shanghai as Huawei's guest, met its researchers, and handled its newest phone on camera. That does not make the physics wrong, but it shapes which questions get asked and which get skipped. No rival architect appears on screen, no skeptical operator is interviewed, and the competing answers from the NVLink camp or the Ethernet camp never get a fair hearing. Viewers should treat the trip as valuable access journalism with a clear host, not as a neutral laboratory comparison.

The practical takeaway survives all these caveats. Whether UnifiedBus wins or not, the industry's bottleneck has moved from raw transistor speed to the cost of moving data, and every buyer of AI compute should now ask for performance per watt instead of peak FLOPS. Operators planning clusters in the next two years can already demand MFU figures, interconnect power budgets, and failure-interval statistics in tenders. If open specifications plus Ethernet compatibility let smaller vendors join in, even skeptics will benefit from cheaper, cooler fabrics no matter whose logo ends up on the rack.

Sources

9 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

artificial intelligence · huawei · unifiedbus · data centers · nvlink · energy

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…