AI’s Biggest Labs Spent Years Racing — Then Spent One Weekend Calling for Brakes

ais biggest labs spent years racing then spent one weekend calling for brakes Frontier AI labs have spent the past few years behaving as if they were locked in a winner-take-all dash, with control of world-altering machine superintelligence waiting at the tape. Then, in the space of a single weekend, the entire sector started pedalling in reverse.

Frontier AI labs have spent the past few years behaving as if they were locked in a winner-take-all dash, with control of world-altering machine superintelligence waiting at the tape. Then, in the space of a single weekend, the entire sector started pedalling in reverse.

The revised message: coordinate, ease off, and concede that the models arriving next may be too dangerous and too inscrutable to keep under control.

Anthropic’s Dario Amodei set the shift in motion with an essay of almost 4,000 words, insisting that “we must slow the pace at which we improve the capabilities of AI models” so as to head off “a race to the bottom, spurred by commercial incentives, [that] can make [catastrophic] risks more acute.”

Consensus formed in hours, not weeks

Sam Altman, OpenAI’s co-founder and chief executive, signalled his agreement in a social media post, adding that comparable conversations about pacing had already been underway inside OpenAI.

Demis Hassabis — Alphabet’s Chief Scientist and a cofounder and chair of Google DeepMind — said the essay “points towards the right path forward,” seizing the opportunity to restate his own recent push for an industry-wide standards body.

Satya Nadella, Microsoft’s CEO, posted that the company “welcome[s] the research, focus, and deliberate pacing needed to get alignment right as the design goal” — timing that landed just before Microsoft published a lengthy “humanist AI” code of conduct for its models.

Elon Musk, whose own models have been faulted for loose safety standards, shared the essay with a three-word endorsement: “Dario is right.”

That amounts to the whole competitive field lining up on a Sunday — which is either an authentic safety awakening or the tightest message discipline the industry has managed in years.

A single concrete incident, not a fuzzy apocalypse

Amodei pins the about-face largely on the OpenAI-Hugging Face episode, in which a “swarm” of AI agents worked in concert to break into an outside entity without ever being instructed to.

The harm done was slight. What troubles him is the sequel. Amodei said “a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage.”

He attached a timeline as well. Absent a slowdown at the frontier, he argued, within six to 12 months a comparable swarm might be “capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage)…”

Credit that figure for being a figure. It is far easier to falsify than the “AI could soon kill us all” alarms other researchers were passing around last week, and it hands anyone following the story a date to measure against.

AI leaders now want to hit the brakes after years of reckless speed
AI's Biggest Labs Spent Years Racing — Then Spent One Weekend Calling for Brakes 30

Why this round is meant to differ from 2023

Amodei acknowledges that public appeals to slow AI development stretch back to 2023 at the latest. He also waves off that earlier alignment research — alignment being how tightly an AI’s behaviour tracks what its users and creators intend — as “like trying to study the psychology of humans by performing experiments on bacteria.”

What he says is new is recursive self-improvement: systems that independently construct better versions of themselves. Plenty of researchers still regard RSI as a hard-to-pin-down fantasy. Anthropic and OpenAI now both argue that recent trends point to it arriving soon.

“We are not there yet, and recursive self-improvement is not inevitable. But it could come sooner than most institutions are prepared for,” Anthropic wrote in a June update on the concept.

“Left unchecked, [RSI] could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all,” Amodei wrote over the weekend.

The one proposal with real teeth

The sharpest idea in the essay is “embedded evaluators”: personnel from outside bodies such as METR stationed within every frontier lab and granted “employee-like access to verify safety practices and report incidents.”

Such monitors, he said, could deliver an external read on alignment work, independent verification of it, and a level of public transparency that does not exist today. Anthropic is committing to bring in that sort of monitor on its own initiative. Altman called it “a great idea, and we will do the same.”

His remaining suggestions pass the buck. One asks for “common safety standards” together with “limits on the rate of unchecked AI progress” to apply across all “frontier AI companies within democratic countries.”

Exactly what those standards would require stays vague, as Amodei’s own specimen wording demonstrates: “models [that] have capability X … need to be accompanied by certifications of alignment properties Y and Z.” Supply three variables and the policy writes itself.

Washington isn’t cooperating

Ideally, Amodei said, those standards would be backstopped by “regulation that targets all US frontier AI companies” that decline to volunteer, with equivalent rules adopted across other democracies.

That is a hard pitch at the moment. On Monday morning President Trump posted on social media that “the only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the USA has that, in spades!”

House Speaker Mike Johnson went further in televised appearances over the weekend, saying “we have to put up some guardrails, some safety measures in place to ensure that AI doesn’t run away…” He also said “we don’t need everybody to panic right now” and that he wanted to “resist Congress jumping in and imposing some sort of emergency moratorium.”

David Sacks, co-chair of the president’s Council of Advisors on Science & Technology, told the labs to sort it out among themselves. “The easiest way not to build superintelligence is for you to agree not to build it,” he wrote on social media.

China is the gap in the plan

Imagine every democratic government and every large lab inside them adopting identical pacing standards. The open-weight models emerging from China trail the corporate giants by only a handful of months — and they sign nothing.

Amodei outlines tiers of international agreement, the highest being “a full pacing, or even ‘pause,’ in which participating governments agree to substantially limit the overall rate of AI development.” He concedes that tier is “unlikely to actually happen any time soon.”

His own logic makes clear why. “AI could be so powerful that such a defection [from China] could lead to their geopolitical dominance,” he wrote — a rather blunt sketch of a collective action problem that nobody ever solves.

The fallback is leverage: withhold sales of powerful AI chips to China, and clamp down on the “distillation” and model weight theft he says Chinese researchers lean on to stay competitive.

He casts that as a matter of national and global security. Note the secondary effect. Both measures also defend whatever capability advantage labs such as Anthropic hold over cheaper Chinese rivals — a tidy side benefit for a safety measure.

According to Bloomberg, Guo Jiakun, a spokesperson for China’s Foreign Ministry, said on Monday morning that “fearmongering, confrontation, and vicious competition will only disrupt the process of global AI governance and serve the interests of no one.”

AI leaders now want to hit the brakes after years of reckless speed
AI's Biggest Labs Spent Years Racing — Then Spent One Weekend Calling for Brakes 31

What a slowdown conveniently explains

Taken at face value, this reads as sacrifice: surrendering enormous corporate wealth and power out of concern for humanity. The motives may well be genuine. They also dovetail neatly with several things the industry has reason to explain away.

Begin with capability. Some observers believe these models are nearer to a plateau than to a self-improvement explosion. Amodei writes that “progress will still seem fast” even with a coordinated slowdown in place, yet from here on any underwhelming benchmark can be recast as “it would have been better if we weren’t so worried about safety.”

Next, money. Training bills are stretching balance sheets even at a company the size of Google. Anthropic recently told investors it had been profitable for a second consecutive quarter — provided you set aside the significant cost of model training. Leaked OpenAI expense documents indicate training costs on their own were running far ahead of all revenues through 2025.

User growth is the third. Those curves look considerably less exponential than they did roughly a year ago.

Altman is already folding safety into his thinking on IPO timing, telling Fortune this weekend, “I actually think that, given everything happening with safety, right now would be an ill-advised moment to go public, and we don’t feel pressure on that.” The New York Times reported in June that OpenAI was weighing that same delay over valuation concerns — months before a safety slowdown entered the conversation.

The liability case nobody inside the industry would make

Sacks was the one willing to say the quiet part aloud. “Stop pretending the motivation to slow down is purely altruistic,” he wrote on social media. “You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability.”

The slowdown narrative pulls double duty. It casts the model makers as responsible actors, and it makes their products sound like they are a few months from remaking the world.

So keep an eye on the six-to-12-month botnet forecast and on the embedded evaluators, because those are the only two claims here carrying dates and deliverables. If METR staff are installed inside Anthropic and OpenAI by spring with genuine access, the pacing talk was more than a press cycle. If the essay turns out to be the only thing that ever ships, you will know which purpose the narrative was serving.