agentic-ai
Anthropic's pace-the-frontier plan is not a pledge. It is an audit-access test.
Amodei's essay proposes embedded auditors and treaty-style agreements, but the only concrete near-term step is letting METR back into pre-deployment testing. Whether that access comes with disclosure rights, not just another pilot, is the part worth watching.
On 12 September, Anthropic CEO Dario Amodei published an essay calling for a controlled slowdown in frontier AI development, warning that recursive self-improvement could threaten the stability of the internet within six to twelve months. His three-step plan proposes embedded auditors inside AI companies, shared safety standards across labs, and global agreements modelled on the SALT disarmament treaties. The concrete commitment attached to the essay is narrower: Anthropic says it will give third-party evaluators, named as METR, access to its models to verify adherence to its own safety practices.
A pledge with no enforcement mechanism
Read as a policy announcement, the essay asks a lot of everyone except Anthropic. Embedded auditors, shared standards and SALT-style treaties all require other labs, or governments, to act. Nothing in the essay commits Anthropic to slow its own model releases, cede a veto to an external body, or accept a defined capability threshold beyond which it would pause. A company that continues shipping frontier models on its existing schedule while calling for slower development elsewhere is not describing a change in its own behaviour.
The part that is checkable
The METR access is different, because it is specific enough to verify. METR has previously said its work with frontier labs, including Anthropic, involved informal pilot arrangements conducted under non-disclosure agreements, with no obligation on either side to publish findings or let them inform deployment decisions. If the access granted now repeats that pattern, the essay changes nothing measurable. If it instead comes with a formal mandate, an obligation to publish, and a real chance to delay a release, it is the first external check with teeth that a frontier lab has granted. Nothing published so far says which one this is.
A safety commitment that a lab can rescind, redefine, or keep under NDA is not a commitment. It is a claim awaiting evidence.
What changes for anyone relying on Anthropic's safety claims
Teams that cite Anthropic's public safety commitments in their own vendor risk assessments should stop treating the essay itself as evidence of anything. The thing to track is whether METR, or any other named evaluator, publishes independent findings on a model before it ships, and whether Anthropic's release timeline has ever been altered by what an evaluator found. Absent that, "we let auditors in" and "we ran the same pilot programme we have run since 2023" are indistinguishable from the outside, and a compliance framework built on the essay's language rather than the access terms is built on the wrong document.
Sources
Written by the Devence Lab research team.