Back to feed

Self-Improving AI Alarm: Why an Anthropic Resignation Shook the Safety Debate

A DW News segment examines warnings from Anthropic insiders that self-improving AI could pose an existential risk within a decade, balanced by a RAND Europe researcher who calls it a serious but unproven future risk, plus the US-China race and EU regulation angles.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — tKQBcdRw21s
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Artificial intelligence now sits inside ordinary routines, from drafting messages to writing code, which is why a stark question lands so hard: could the same systems one day threaten humanity itself. A DW News segment takes that question to researchers and policy voices after a week of alarming posts from inside the AI industry. The tone is urgent but genuinely inquisitive, asking how seriously each claim should be taken.

The trigger was the resignation of Jacob Coxon, who spent about three years on pre-training research at OpenAI and Anthropic before quitting and posting that frontier labs are gambling with our lives. His message holds that people building these systems privately expect they might kill us all before the decade ends, while public statements stay polished for the cameras. He frames the danger less in today chatbots than in self-improving superintelligence that could outpace every check placed on it.

This exit might have faded as one departure story had Evan Hubinger, who leads the Alignment Science team at Anthropic, not publicly backed it. Hubinger wrote that staff sincerely believe the technology could wipe out humanity, putting his own estimate above ten percent within ten years, while stressing that present models pose little such danger. His widely shared message added an uncomfortable admission: the lab has no solved plan for keeping superintelligence aligned and is not clearly on course to one.

At the heart of the dispute sits recursive self-improvement, the idea that an AI system could help design its own successor, from code to research methods to compute use. Each generation built with machine help could arrive faster and less transparently than the last, forming a feedback loop that testing and governance cannot keep up with. That loop, rather than any single model release, is what the warners say compresses the schedule from decades to years.

The program guest from RAND Europe, a director working on emerging technologies and resilience, calls the worry legitimate but draws a firm line. In her account, no fully autonomous system has yet demonstrated the complete cycle of improving itself into something qualitatively more capable. Pieces of the hypothesis are visible, though, in agents that wield tools, pursue extended tasks, and sometimes act in unintended ways. So the honest label is serious future risk, not established fact about deployed systems.

Asked how extinction could concretely happen, she points to loss of control: a highly capable agent with deep access to tools and infrastructure, pursuing a poorly specified goal that humans cannot reverse. Reaching civilization-scale harm would require stacked failures, extreme capability plus real-world access plus weak safeguards plus missed detection plus slow institutional response. Each layer is individually plausible enough to plan against, and none of them requires assuming catastrophe is inevitable.

Politics answers in a different register. Asked about extinction, the US president deflects toward competition, asserting roughly a one-year American lead over China, a margin specialists describe as a few months at most. The RAND guest widens the frame beyond two capitals to firms, supply chains, and universities, warning that rivalry rewards speed while treating safety work as a cost. When deployment velocity becomes the scoreboard, restraint and transparency lose by design.

The segment closes on credibility and remedies. The Elon Musk camp dismisses the alarm as a sophisticated maneuver to invite regulation, while the RAND researcher replies that motives deserve scrutiny but claims should be judged on evidence and on whether safeguards actually improve. Her concrete pointer is Europe: the EU AI Act enforcement arm, the AI Office, with new powers to inspect models and penalize providers. The interview ends where it began, with the anchor noting this debate will return, because the underlying race has not slowed.

Visualization: nodesdaily AI

AI commentary

"What struck me most is the split screen: the people closest to the technology sound the most frightened, while the political response is still framed as a race to be won. I find the RAND researcher caution the most useful voice here, because it separates a serious future risk from an established fact. My read: plan for loss of control, but judge every probability claim by who measured it and how."

AI assessment

The strongest counter comes from Timnit Gebru, who argues in a WIRED interview that extinction talk distracts from harms already here, such as autonomous weapons, and that it can serve the commercial interests of the labs issuing the warnings. I take this steelman seriously: fear of hypothetical superintelligence can concentrate attention and funding on a few frontier firms while diffuse, present-day damage gets less scrutiny. The right test is whether a warning comes with verifiable safeguards, not only with a striking number.

What the segment does not test is the evidence base behind the numbers. A greater-than-ten-percent chance of human extinction within a decade is a personal credence, not a measured frequency, and the August alignment report from Anthropic rates catastrophic risk from current models as low while flagging harder-to-detect behavior in future systems. I also missed two pieces of context the press later added: a summer string of incidents in which autonomous agents carried out cyber-attacks, and a report that Anthropic withheld its newest model from the UK safety institute. Both cut in opposite directions, which is exactly why single-number claims deserve a second source before they travel.

On interests and checkability, I note that Anthropic built its brand on safety-mindedness while competing in the very race its researchers criticize, and its founder publicly favors regulation, so scrutiny of motives is fair. But motive scrutiny is not refutation, and the verifiable core stands: the alignment team openly states it has no solved plan for superintelligence alignment, and lawmakers in Washington and London are already converting the warnings into treaty and oversight proposals. I treat the corporate framing as suspect and the technical admission as data.

My practical take is tiered. If you build or deploy autonomous agents, harden tool access and logging now, because loss-of-control scenarios all begin with broad permissions, not with superintelligence. If you regulate or invest, watch the EU AI Office new enforcement powers and the multinational treaty letter rather than social-media probability posts. And if you simply use these tools, keep the RAND distinction in mind: a future risk worth planning for is not the same as a present capability, and confusing the two leads to both panic and complacency.

Sources

8 links; 3 of them also cited by 4 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

anthropic · ai safety · superintelligence · eu ai act · us china ai race

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…