Back to feed

The Next Trillion-Dollar Chip Race Will Be Won by Selling Systems, Not Chips

Inference spending has overtaken training, close to two percent of United States output is flowing to artificial intelligence infrastructure, and rack-scale integrated systems are overtaking part picking. Austin Lyons explains through the prefill versus decode split and the rise of neoclouds where the next trillion-dollar chip company could emerge.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — bXNJU3RhW1I
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

No company has yet taken a chip built specifically for large language models all the way to market. On Celestia Capital's Tech Surge podcast, semiconductor analyst Austin Lyons explains it as pure market logic: if the benefit is unclear, no one changes the factory line. The numbers are now telling a different story. Close to two percent of United States gross domestic product will go to artificial intelligence infrastructure this year, almost twice the 2025 level. The quiet break happened last year, when spending to run models in production overtook spending to train them, and that crossover is reshaping every link from buyer to designer to supplier.

How cloud buyers purchase hardware has flipped. Lyons recalls that a decade ago clouds pushed for commoditization through standards bodies, bargaining part by part to drive prices down. Today the pull is opposite: customers prefer a fully integrated rack from one supplier rather than assembling parts themselves. The reason is pace and simplicity. In a boom, the winner is not the cheapest bill of materials but the fastest system to stand up and put into production.

The case for selling systems starts with model scale. A frontier model with more than two trillion parameters still needs one to several terabytes of weight storage even after quantization. A single accelerator, however, holds only about 288 gigabytes of high bandwidth memory. That forces the model to be sharded across four or eight accelerators, and once sharded, thinking in racks becomes unavoidable. Nvidia's Grace Blackwell NVL72 embodies the answer: 72 graphics processors and 36 central processors working as one system with networking, power and cooling.

Nvidia's edge is that it got to rack scale first. According to Lyons, the company stitched together the full design, the software stack and the compiler toolchain and made the complex rack turnkey. A model team can write code and run it, with rack complexity hidden. The industry always wants competition and commoditization, yet when integration pain is this high, the supplier who delivers a working rack first sets the terms.

Why would a cloud buyer accept high margins at every layer instead of mixing vendors? The answer is time to market. Lyons points to Elon Musk and xAI standing up data centers in record time on integrated racks, where speed to useful tokens mattered more than squeezing cost. AMD's Helios rack shows the same pull. Everyone knows a mix-and-match build could be cheaper on paper, but when deployment windows are measured in days, an integrated rack is worth the margin.

That raises the bar for startups. Lyons notes that what once took a few million dollars in proof-of-concept rounds now takes hundreds of millions for a chip startup aiming at rack-level product. Incumbents compound the advantage with mergers such as Mellanox and decades of system know-how. Yet the window has not shut. As interfaces mature, narrow niches where a startup can live without cloning an entire rack begin to open, especially where the workload has a distinct bottleneck.

The inference workload itself is splitting, and the split shapes silicon choice. In the conversation the two phases are described as prefill and decode. Prefill is the parallel read of a long prompt, hungry for compute. Decode is the sequential generation of tokens, dominated by how quickly data can be moved, which makes memory bandwidth the decisive factor. Natural engineering says to run prefill on compute-heavy pools and decode on memory-centric pools, and to disaggregate them when it pays.

Nvidia codified the split in software with an orchestration layer called Dynamo that manages prefill and decode pools across the same graphics processor fleet. Interestingly, earlier bets by Groq and Cerebras now make sense. Both built architectures around abundant on-chip static random-access memory instead of capacitor-based dynamic memory. That choice gives them very high bandwidth on chip, a near perfect fit for the decode side where moving data fast matters most.

The binding constraint in data centers is power. Once a campus is granted one hundred megawatts, the business question becomes how to earn the most revenue per megawatt. Lyons frames design choices as finance questions under a fixed power, space and cooling envelope: which silicon mix yields the most inference output. The math even lets former bitcoin mining sites reinvent themselves as cloud capacity where cheap power and asset-backed capital let them move faster than slow-building hyperscalers.

A new intermediate layer called neocloud grew from that gap. The conversation explains why even Microsoft rents graphics processors from neoclouds while building its own campuses: power is scarce, off-balance-sheet renting is flexible, and demand calendars are unforgiving. Neoclouds bring experience financing assets and stand up capacity quickly, hyperscalers use that rented capacity as a buffer while expanding owned campuses, and the result is an intermediate tier now worth more than one hundred billion dollars.

Whether neoclouds can be durable is the open question. They serve hyperscale customers while competing with them indirectly. Lyons expects consolidation but not a simple acquisition by chip vendors, partly due to regulatory friction. Neoclouds can instead aggregate diverse silicon themselves and hide complexity by selling tokens as a service, matching each job to the right hardware. If that works, it creates a route to market for novel silicon, because the cloud layer willing to shoulder heterogeneity lowers adoption cost for customers.

Will clouds end up with one chip type or a wide portfolio? Lyons points to central processors as precedent: no cloud offers just one central processor today, and the same logic will apply to inference. Picture a curve where the horizontal axis is interactivity as tokens per second per user and the vertical axis is total throughput. Each point on the curve is a different product, from low-latency chat to batched summarization to cheapest small-model serving, and each point is cheapest on a different hardware blend.

The trillion-dollar opening may sit in the middle of that curve, not at the frontier. The discussion is fixated on the largest two-trillion-parameter model, yet the largest volume may be in seventy-billion-parameter open models. Lyons observes that older Hopper and Ampere cards remain competitive there, so a new accelerator that targets that crowded middle with a memory-first design can capture real volume at a better price without needing the highest intelligence.

The build-versus-buy dilemma for silicon is also clarifying. Leading model providers buy merchant graphics processors at scale while co-designing custom chips with partners such as Broadcom and Marvell, with chip and software teams working together on specific workloads. Merchant vendors, meanwhile, invent not for one customer but for the broad market, letting architecture ideas drive the roadmap. As with OpenAI, both paths advance in parallel rather than excluding each other.

Chip design itself is accelerating with assistance from artificial intelligence. The scenario sketched on the show is striking: a design that once needed three years and a hundred engineers could be done in one year with the same or a smaller team, cutting cost to a third. A two-thousand-dollar addition to the bill of materials that a finance chief would veto becomes viable at seven hundred dollars. As the barrier falls, teams that never considered a custom chip begin to find the math workable.

The same acceleration helps merchant silicon houses. An experienced engineer uses acceleration tools more effectively than a first-time team, so incumbents can proliferate variants under the same roof. Markets once deemed too small to justify a product line can cross the internal rate of return once cost is halved, and companies start adding semi-custom offerings to catalogs. Niches multiply while the overall market gets more tailored choices.

For startups the picture is two-sided. Competing head-to-head at rack scale is difficult, yet lower design cost opens new variant spaces. Even large vendors will now green-light product lines they once shelved as uneconomic. Competition shifts from one best chip to who can place the right silicon in the right phase at the right price with orchestration that makes heterogeneity invisible to the customer.

The final line of the conversation is the best summary: so much remains unsettled, and the unsettled state is not slowing spending at all. Where power limits, memory architecture and software orchestration intersect, who will earn the trillion-dollar crown stays uncertain. Yet the rise of inference work, the normalization of rack-scale systems and the capital model of neoclouds already look entrenched. In this race, the winner will not be the one who waits for clarity, but the one who builds the system under uncertainty and maximizes revenue per megawatt.

Visualization: nodesdaily AI

AI commentary

"What struck me most is how the chip stops being the lone hero and the model that sells a system at the intersection of power, memory and software orchestration steps forward. Even through a hardware lens, the finance question has become decisive."

AI assessment

The strength is the reframing of artificial intelligence infrastructure as a contest of systems and business models, not just chips. Placing the prefill versus decode split on compute and memory axes, explaining neocloud financing through power and balance-sheet logic, and tying rack scale to a finance question about revenue per megawatt gives the view a consistent spine that also supports the portfolio and variant-multiplication arguments.

Limits show in how headline numbers are presented and dated. Claims such as two percent of gross domestic product or inference spend overtaking training for the first time arrive rounded and from a single conversational source, with unclear scope on whether they are annual or quarterly and which cost buckets are included. The promised four conditions for the next trillion-dollar chip company are announced up front but never unpacked as a testable list, leaving the threshold vague.

The pragmatic takeaway is to buy integrated racks over part picking, prioritize revenue per unit of power, and plan inference not on one chip family but on a portfolio. Keeping older generation cards for the middle tier around seventy billion parameters, backing the decode phase with memory-centric accelerators, and orchestrating separate prefill and decode pools with a layer like Dynamo improves the cost and speed balance over the next year.

For startups and cloud teams the practical path is to define the niche around a memory-sensitive bottleneck rather than challenging rack scale head-on. The bitcoin-miner turned neocloud example shows that where power and capital access exist, speed becomes an edge, and running merchant and custom silicon in parallel spreads risk. Even as artificial intelligence assistance lowers design cost, tape-out and portfolio discipline remain essential, or speed gains simply inflate variant sprawl.

Sources

6 links; 1 of them also cited by 2 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

ai infrastructure · chip · inference · rack system · nvl72 · neocloud · memory

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…