Back to feed

Google is selling its own TPUs: the eleven-year closed door finally opens

In 2026 Alphabet said it would begin delivering TPU accelerator systems into customer-owned data centers and opened parts of its software stack to run on other hardware. The move turns a defensive advantage into a standard, and the industry's real bottleneck is neither design nor capital but high-bandwidth memory and grid power.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — HneK5AIUUDo
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Not on September 15, but on April 29 , Alphabet said something unexpected on its quarterly earnings call. CEO Sundar Pichai addressed a broad range of buyers, from AI labs to capital markets firms, with a single sentence: as TPU demand grows, the company will begin delivering TPUs in hardware configurations into the data centers of a select group of customers. For years the only path had been to rent capacity through Google Cloud, and the silicon itself never left Google's buildings. Pichai framed the change as a way to expand the addressable market. On the second-quarter call CFO Anat Ashkenazi confirmed that Alphabet recognized revenue from TPU system sales delivered to customer data centers for the first time, with the large majority of contracted revenue expected to land in 2027. She pointed to a project with a Blackstone-backed venture as the example. As reported by DataCenterDynamics, these hardware agreements sit inside Google Cloud's $462 billion backlog, with only a small percentage expected as revenue late in 2026 and the majority converting in 2027.

The video describes a three-part announcement . The first two are substantially real; the third is largely invented. The first is correct: TPU systems now ship out of Google's own buildings and get installed in customer facilities. The second points in the right direction: Google is running TorchTPU in preview with select customers, giving PyTorch native support on TPU, and expanding vLLM coverage. But the third part, described in the video as a buried clause on page nine — that Gemini models can run on customer-owned infrastructure outside Google Cloud — appears in none of the official disclosures. What was announced was the sale of a chip, not the relocation of a model.

The timing question has a documented answer, and it is the video's strongest section. Google's decision to design its own chip in 2013 marks the beginning of enterprise AI. The story told on the Google Cloud blog is this: chief scientist Jeff Dean realized that if every user in the world spoke to Google for only three minutes a day, the company's data center count would have to double. At the price of compute in that era, that was a bill nobody could pay. Designing an ASIC normally took years; the TPU went from first design to production deployment in fifteen months . Jouppi's 2023 retrospective explains why that was not a fluke: a single-purpose schedule, a 700 MHz clock rate that made timing closure easy, and a mature 28 nanometer process.

The TPU architecture is the technical core of the story. What engineers call a systolic array is a grid of thousands of small multiplier cells wired directly to each other, with data flowing like water through a pipe and being multiplied at every junction without returning to a central controller. That design removes the round trips to memory that general-purpose processors, and to a lesser degree GPUs, require for neural network work. The TPU v1 paper published on arXiv in 2017 measured the chip at 80 times the performance per watt of contemporary CPUs and 30 times that of GPUs . What the video omits is a detail that matters: the first TPU performed inference only , while training stayed on GPUs. Google later split the line, and by the eighth generation TPU 8t handles training while TPU 8i handles inference.

Nvidia's most misunderstood advantage is software, not silicon — and on this point the video's central claim holds up. Building your own chip protects only yourself; it does not change the industry. For fifteen years rivals built impressive chips and accomplished nothing, because the CUDA software layer was written into the reflexes of every tutorial, every library and every graduate student. SemiAnalysis estimates that the total cost of ownership for deployed TPU systems runs roughly 30 to 44 percent below comparable Nvidia GB200 and GB300 deployments. That is the gap Google is attacking: TorchTPU is an attempt to hand the software's advantage to people who are not renting Google's hardware to get it.

The competitive picture shows an uncomfortable symmetry . Amazon has shipped Trainium and Inferentia for years and prices them seriously. Microsoft runs a strong in-house accelerator program. Meta designs MTIA accelerators with Broadcom for ranking and recommendation workloads and plans mass deployment of MTIA 500 by late 2027. Google is not only pressuring Nvidia. The Korean market shows it. S&P Global Ratings expects South Korea to invest roughly 1,200 trillion won in data center construction through 2035, which would make it the second-largest market in Asia-Pacific after China. As reported by Seoul Economic Daily (SEdaily), Google met National AI Strategy Committee vice-chair Ha Jung-woo on September 9 to discuss priority TPU supply, university infrastructure and on-premises installation; analysts read the meeting as a signal that Korea is becoming one of Google's main TPU hubs.

On the numbers, the analyst estimates point in the same direction as the claim, even if the scale does not. Morgan Stanley projects direct TPU sales of $84 billion in 2027 and $108 billion in 2028 , consistent with TF International Securities' assumption of 0.3 gigawatts shipped in the second half of 2026, 3.2 gigawatts in 2027 and 4.2 gigawatts in 2028. TF also models the business at a 30 percent gross margin, which is an admission that hardware trading is structurally thinner than cloud rental. Google Cloud chief Thomas Kurian said the TPU business is more than twice the size of the next-largest hyperscaler accelerator business, and that server payback is under two years, with Google's own processor paying back in about half the time of a GPU. Those are management estimates, not audited quarterly data. The IEEE ComSoc technology blog pulled this picture together: Kurian's 'more than twice the next hyperscaler' remark at the Goldman Sachs investor conference sits in the same file as Morgan Stanley's revised $84/$108 billion outlook.

The cash figures the video leans on hardest are also mostly wrong. The $93 billion number for 2026 capital spending conflates a quarterly figure with an annual one. The real figure is a $180 to $190 billion capital expenditure guidance for 2026, raised from a prior $175 to $185 billion range. Second-quarter capital spending reached $44.9 billion against $39.1 billion of operating cash flow, leaving negative $5.9 billion in free cash flow . One analyst account goes further, reporting that Alphabet sold $80 billion of new equity for the first time since its 2004 IPO. Against that, Alphabet shares fell about 3.8 percent on a day Google Cloud revenue grew 82 percent year over year.

Nvidia's side is more complicated. The company reported $215.9 billion of fiscal 2026 revenue , up 65 percent, with record quarterly data center revenue of $62.3 billion. Gross margin came in at 71.1 percent for the full year and 75.0 percent for the fourth quarter — so the widely cited 70 percent figure was roughly right by accident. The stock, however, has retreated from a $236 record set on September 14, slipping to $225.51 on September 23 and around $221 on September 24, leaving it roughly 15 percent below its peak . According to price action reported by FXLeaders, the shares slipped to $225.51 on September 23 and toward $221 on September 24, with $230 watched as resistance and $220 as support. The video's framing of a first real problem in three years points the right way: TechCrunch's reading is that what the market is repricing is not weak demand but the price of compute itself.

To understand the move, it helps to know why the chip was never sold for eleven years. In 2016 AlphaGo beat the world's best Go player in a live broadcast, and hundreds of reports credited the victory to algorithms and neural networks. Very few people asked what the machine was running on. It was running on Google's own silicon, and the appearance was a demonstration, not a launch. The logic was sound at the time: there is no reason to hand a competitor the one thing they cannot buy. Every enterprise conversation Google Cloud had since 2018 rested on the idea that only Google had this. That background has now changed: because Google gives the software away, selling the hardware carries no competitive downside.

Two bottlenecks: memory and power

The video's most concrete point lands here: the real limit on AI hardware is not design but high-bandwidth memory. HBM stacks thin silicon layers on top of each other and drills thousands of vertical channels through them, connected at a precision measured in fractions of a micron. If a single layer fails, the entire stack is worthless. Reporting from NetworkWorld notes that Samsung, having shifted capacity toward HBM and premium DRAM, raised the price of a 32 GB DDR5 module from $149 to $239 in 2026, a 60 percent increase. SK Hynix said its HBM, DRAM and NAND capacity was "essentially sold out" for 2026. Micron exited the consumer memory market entirely. Per a Micron executive, HBM consumes roughly three times the wafer capacity of standard DRAM per gigabyte.

The scale of that constraint is worth stating plainly. Nvidia has secured commitments covering roughly 37 percent of global HBM output for 2027 , and in the quarterly filing disclosed in August its supply and capacity commitments jumped from $119 billion to $279 billion in a single quarter. Morgan Stanley estimates that Nvidia, Alphabet and AMD together account for about 85 percent of projected 2027 HBM production , leaving 15 percent for every other AI hardware developer. Independent estimates put unit shipments rising from 2.76 million in 2024 to as many as 8.8 million by 2027. Allocation, not design, has become the competition. Micron's target of 100,000 HBM wafers per month by year end does not change that: industry estimates put Samsung and SK Hynix at 150,000 to 200,000 wafers each. According to a CSIS analysis, this is a structural shift rather than a cyclical swing: data centers will consume about 70 percent of global memory output in 2026, with meaningful new capacity arriving no earlier than 2027-2028.

The second bottleneck is power and grid interconnection , and it may be the harder one. Accuris tracking data puts US data center energy demand at 80 gigawatts in 2025, projected to reach 150 gigawatts by 2028, with global data center electricity consumption approaching 1,050 terawatt-hours — roughly the size of a mid-sized national grid. Data centers are set to consume 70 percent of memory production, and the same report says 30 to 50 percent of planned 2026 AI data center capacity is slipping to 2028 because of grid interconnection queues and construction bottlenecks. Local commissions in Virginia, Georgia and Alabama are still weighing whether to let data centers connect while residential rates rise. So the video's closing instinct is right: the binding constraint on this race may not be chips at all, but permission to use land and connect to the grid.

Five signals worth watching

The video's most useful section is its watch list, even though the figures are wrong. The first signal is a third party publicly confirming production workloads running on this hardware in its own facility, because announcements alone have a poor track record of surviving contact with delivery schedules. The second is a customer publishing its cost per unit of work against what it used to pay. When that number appears, an actual economic comparison replaces the marketing. The third is the incumbent's response: cut prices, bundle more aggressively, or race the architecture. A swift price cut signals distress; silence signals confidence. The fourth signal is how Amazon and Microsoft position themselves. Both run their own silicon, which makes Google's chip a competitor. But Google's software layer is free and runs on anything, including their own. Adopting a rival's open package is a hard choice without government cover. The fifth is long-term supply agreements: whoever secures HBM capacity through 2029 has already won fights that have not started. Meanwhile Morgan Stanley supplies a concrete marker — Anthropic's plan to expand TPU deployment from 1 gigawatt this year to 5 gigawatts next year should be read as evidence of supply tightness, not of demand weakness.

Where the video goes wrong is not the existence of the announcement but its date and its scale. What actually happened is interesting enough: a chip that was refused for sale in 2015 opened up in the middle of the largest infrastructure shift in a generation, and it did so by breaking its own margin model to become a standard. To see how bold that is, consider who else could do it. The endgame of competition is not outbuilding a rival; it is becoming something a rival cannot work around. After eleven years, the preference to keep the old tool inside the walls is gone.

Visualization: nodesdaily AI

Key moments

  1. The 300 billion dollar claim
  2. A three-part announcement
  3. Where the capex number comes from
  4. GPU versus TPU
  5. Nvidia's software moat
  6. The Broadcom and TSMC chain
  7. The 2013 voice search calculation
  8. Fifteen months to silicon
  9. AlphaGo and the secrecy decision
  10. The memory bottleneck
  11. Grid interconnection queues
  12. The September 2026 framing
  13. Positioning it for inference
  14. The page nine claim
  15. The 2017 passenger analogy
  16. Five signals to watch
  17. Hardware or platform

AI commentary

"The video's framing is compelling but most of its numbers do not survive contact with the record. There is no September 15 announcement and no 1.4 million chip figure; the real disclosure came on Alphabet's April 29 earnings call, and the first hardware revenue was recognized in the second quarter. Even so, the underlying thesis holds: Google is deliberately trading its own hardware margin to become the default platform."

AI assessment

The video's strongest claim is that the software ecosystem, not the hardware, is what makes Nvidia durable, and it sets that argument up correctly. For fifteen years rivals built their own accelerators and none of it worked, because the obstacle was habit, not architecture. Google's TorchTPU push attacks exactly that chain. But the numbers offered in support are largely invented. The SemiAnalysis estimate of a 30 to 44 percent cost advantage is a single source and applies to deployed systems, not a chip-to-chip comparison. Higher throughput is not automatically lower total cost, because software porting effort and capacity planning do not appear in that figure.

The most serious problem is the fabrication of the date and the scale. The claimed September 15, 2026 date is in fact the April 29 earnings call. The claim that Nvidia fell more than 6 percent across two sessions, erasing $300 billion, matches no record: Nvidia was reporting record revenue at the time and its stock was approaching record highs. The 1.4 million unit figure also does not come from Google; it is an analyst estimate that never appeared in the announcement. The Morgan Stanley forecast that does use 1.4 million units for 2027-2028 looks reasonable alongside its $84 billion revenue projection. The video takes a real development and tells it with a wrong date and invented scale, which costs the reader trust.

Most of the speaker's inferences are reasonable, though one deserves more skepticism. The line that Google is not launching a product but opening an exit is elegant but technically vague: Google is shipping physical systems into customer facilities while remaining the cloud provider, and TF International models the business at a 30 percent gross margin, far below the cloud business. That is not walking away from the moat; it is adding a low-margin line of business. Bitcoin's 2017 transition followed a similar pattern, as exchanges opened their own infrastructure and then remained both venue and infrastructure provider for a while. The 30 percent margin figure deserves independent verification: as long as Alphabet declines to break out this business separately, the margin foundation of the whole growth story stays unverified.

The practical implication is clear. This is a directional signal, not a trading signal. Nvidia's results remain strong: revenue up 65 percent, record data center quarters, gross margin at 71 percent. But the valuation now rests less on next quarter and more on who sets the price of compute. TechCrunch, citing Ornn data, points at the mechanism: the spot price of an hour of H100 compute peaked around $3.20 in May and has since declined. When the price of an hour of computing falls, the margin compresses for everyone renting these chips. Three things are worth tracking: disclosed cost per unit of work, the pace at which customers like Anthropic and Crux bring capacity online, and whether Nvidia expands its own HBM allocations in the next quarter.

Sources

11 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

tpu · nvidia · ai chips · data center · hbm · alphabet

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…