Anthropic’s Dario Amodei Calls for the AI Race to Ease Off the Gas

anthropics dario amodei calls for the ai race to ease off the gas Two things have always happened simultaneously at Anthropic: shipping frontier models, and sounding the alarm about them. This week its chief executive put almost all of his weight on the second.

Two things have always happened simultaneously at Anthropic: shipping frontier models, and sounding the alarm about them. This week its chief executive put almost all of his weight on the second.

Dario Amodei argued that the moment has arrived for AI development to decelerate, publishing a sprawling essay that sketches a three-part plan to “pace the frontier.” Translated out of the jargon, the idea is to ease the tempo of training and development so that labs have room to put safeguards in place and regulators have room to assess what is actually being released.

Only the first step is his to make

Anthropic intends to open up broad access to its models for outside evaluators such as METR, as a way of demonstrating its “adherence to safety practices and commitments.” According to Amodei, that is step one — and the company is doing it right now, on its own initiative.

That commitment is the one to keep an eye on, since it is the single item on the list that requires no competing lab and no government to sign off on anything beforehand.

The other two steps need everyone else to cooperate

The second step calls for the sector to move collectively, most likely in tandem with government agencies, to “establish common safety standards as well as limits on the rate of unchecked AI progress.” Amodei limited the scope of that proposal to AI firms based in democratic countries. Since drafting legislation and standing up regulatory machinery is slow work, he suggested companies collaborate on safety standards in the interim.

Put plainly: firms would impose their own speed limits until the real rulebook arrives.

The third step is, in his own telling, the toughest of the lot. It would require authoritarian governments — China and Russia among them — to agree to a slower pace and sign on to a worldwide set of AI safety standards.

He added that it remains vital for the United States and fellow democracies to hold a technological edge over China and other authoritarian states, achieved by restricting their supply of high-powered chips and clamping down on tactics such as distillation, in which a firm trains its AI to mimic the behavior of a stronger model and closes the gap quickly.

The upshot is a pair of instructions — decelerate, yet remain in front — that sit awkwardly together, and the essay makes no attempt to disguise the tension.

What’s driving the alarm

Two developments underpin Amodei’s concern. One is recursive self-improvement, known as RSI, in which AI systems build the following generation of AI and capability gains steepen dramatically. “Left unchecked, it could outrun our ability to understand and control these systems,” he said.

The other is the OpenAI and Hugging Face episode from this summer, in which “a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance.”

One caveat is worth holding onto here: Anthropic’s own Claude sat at the center of a run of rogue AI hacking incidents that have lately drawn scrutiny onto the company.