Back to feed

Nikesh Arora on Where AI Goes Next: Agents, Inference Bills, and Rebuilding From Scratch

Palo Alto Networks CEO Nikesh Arora argues AI convictions now expire every few months, with agents and inference costs reshaping enterprise strategy. He makes the case for rebuilding products from scratch, securing agents at runtime, and funding the data-center buildout without wasting equity.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — fLy3fp63xv4
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Convictions about AI now expire every three to four months, and that churn favors builders over forecasters. Nikesh Arora argues ChatGPT won by shipping at roughly 70% accuracy while incumbents waited for perfection near 100%. Speed created the data flywheel, and the flywheel then compounded into agents, open weights, and new inference patterns.

Agents change the unit of work from a single answer to an outcome pursued across tools and time. Arora describes long-running helpers that plan, browse, call APIs, and revise their own steps. Open-weight models add leverage because firms can inspect, distill, and host them closer to sensitive data. The shift moves value from model novelty toward orchestration, memory, and trust.

That autonomy has a meter running underneath it, which generalcompute illustrates with hours-long runs making around 600 language-model calls and carrying 50k-token context windows. Cost compounds because every retry drags history forward, so teams need checkpointing, scoped memory, and budgets per task. Arora frames this as the hidden tax on useful agents: longer horizons multiply both capability and spend.

Inference now dominates enterprise AI budgets, a turn zylos calls the Inference Flip in early 2026 with inference near 85% of spending. Agentic workloads can multiply token use by 100 to 1000 times versus single prompts, while Gartner sees 5 to 30 times more tokens per task. With H100 rentals near $2.01 per hour and Blackwell hardware roughly 3 times cheaper per token, architecture choices decide margins.

Boards keep raising AI budgets even as deployment lags, a tension kpmg quantifies with average spend near $207M and 54% of firms deploying agents versus 12% in 2024. Yet 65% struggle to scale pilots and 62% cite skills gaps. Arora reads this as an execution crisis rather than a model crisis: money buys experiments, but only workflow redesign and talent turn pilots into production returns.

Why data centers need new money

Hyperscalers enjoy roughly a 1.5 times cost advantage through power contracts, chips, and utilization, while equity-funded data centers look like the worst way to finance concrete and electrons. Arora expects a neocloud and necaler layer to absorb spillover demand, then normalization within five to seven years as capacity catches up. Until then, scarcity pricing rewards owners of power and cooling.

Given $1B, Arora would back AI transformation help and rebuilding products from scratch rather than bolting copilots onto old suites. The logic is that incumbents protect workflows while newcomers redefine them, so retrofits inherit debt. A clean-sheet rebuild lets teams choose agent-native workflows , persistent memory , and policy guardrails as primitives instead of patches applied after launch.

The car analogy makes the rebuild point vivid: Waymo rebuilt driving as driverless, Tesla delivers 60 to 80% self-driving with hands still near the wheel, and Mercedes gets mocked for AI-washed assistants that resemble chatbots on wheels. Customers feel the difference between autonomy and assistance. Arora warns many enterprises are buying the chatbot version while claiming the driverless prize.

Where incumbents and startups each win

Most firms mistake an API core for a real business, then discover distribution and trust decide survival. Drawing on nfx research on disruptive versus sustaining plays, Arora urges a greenfield strategy for new categories instead of defending old margins. With 80,000 customers demanding consistency, Palo Alto cannot experiment recklessly, so it watches about 40 agentic security startups as a live lab before acquiring with humility.

That lab logic shows in paloaltonetworks product moves such as Prisma AIRS 3.0 to discover, assess, and protect AI assets, plus an AI Agent Gateway and the Koi acquisition for agent identity. Discovery maps shadow models and tools, assessment scores risk, and protection enforces runtime policy. The message is that securing agents needs continuous inventory , risk scoring , and runtime enforcement , not annual audits.

Offense is accelerating too, which fortune coverage captures through the Mythos red-teaming test against a roughly $300B company after June export controls, a moment when chief executives now ask to see attacks live. Models find flaws faster than human teams, and Arora says dwell time can shrink from four days toward one minute. Defense must therefore automate triage while keeping humans accountable for edge calls.

Productivity, scale, and the road ahead

Engineering output is rising, yet Palo Alto keeps about 6,000 engineers because demand and backlog keep growing. Coding helpers clear routine work, freeing seniors for architecture, review, and incident response. Arora treats productivity gains as throughput expansion rather than headcount reduction, with backlog absorption and quality review determining whether speed becomes reliable software or merely faster drafts.

Scale funds the strategy, as prnewswire results show fiscal fourth-quarter 2026 revenue of $3.41B up 34% and next-generation security annual recurring revenue of $9.10B up 63%, with a $20B fiscal 2030 target. Cash flow from platform consolidation pays for AI security bets. Investors get a rare mix of growth and discipline, which explains why Arora can promise aggressive acquisition without losing operating focus.

Arora closes with practical advice for chief executives: master the day job first, then engage openly with policymakers including the White House on power, safety, and competitiveness. Dialogue beats distance when infrastructure rules get written. Firms that pair internal transformation with external credibility will navigate the next conviction shift faster than rivals chasing every model release.

Visualization: nodesdaily AI
AreaSignal
Agents cost600 calls, 50k context; tokens up 100-1000x per zylos
Enterprise gapkpmg: $207M spend, 54% on agents, 65% stall
Palo Altoprnewswire: $3.41B revenue, $9.10B NGS ARR

Key moments

  1. Convictions expire in months
  2. Ship at 70 percent
  3. Long agents, large bills
  4. Waymo versus chatbot cars
  5. Forty startups as a lab
  6. Day job plus White House

AI commentary

"Arora is most persuasive when he links autonomy to economics: longer agents mean larger inference bills, so governance becomes a cost control. I find the rebuild-from-scratch argument stronger than the financing forecasts, which depend on power markets few can predict. The practical test is whether security teams can move from demos to measured production control."

AI assessment

Optimists see agents compounding into leverage, while skeptics see compounding bills and brittle autonomy. The generalcompute warning about 600 calls and 50k-token context plus the zylos math of 100 to 1000 times token growth and H100 pricing near $2.01 per hour suggest budgets need guardrails now. The kpmg finding of $207M average spend with 65% struggling to scale and 62% short on skills backs Arora: transformation, not model choice, is the bottleneck.

On industry structure, the nfx greenfield lesson and paloaltonetworks moves like Prisma AIRS 3.0 with discover, assess, and protect plus Agent Gateway and Koi argue for building security around agents from day one. The prnewswire print of $3.41B quarterly revenue and $9.10B next-generation ARR shows platform cash can fund that rebuild toward a $20B goal. Still, hyperscale cost advantages and five-to-seven-year normalization mean timing and power access matter as much as code.

On risk, the fortune account of the Mythos test on a $300B company after June export controls, with chief executives newly eager for live demos, supports faster red-teaming paired with faster patching. Shrinking exposure from four days toward one minute only helps if fixes ship at machine speed. Readers should therefore demand measured pilot economics, verified runtime controls, and clear human ownership before declaring any agent deployment truly driverless.

Sources

8 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

nikesh arora · ai agents · inference costs · ai security · palo alto

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…