Anthropic has proposed embedded third-party evaluators to monitor AI safety commitments, while OpenAI has indicated it would adopt the same arrangement. The broader debate also includes calls to slow AI development and demands for open-weight models.
What Dario Amodei proposed
Anthropic CEO Dario Amodei called for slowing the pace at which AI model capabilities improve and urged leading AI companies in democratic countries to coordinate on common safety standards and limits on unchecked progress. As a first step, he proposed embedding third-party evaluators, such as METR, and Anthropic committed to the arrangement.
OpenAI follows the evaluator strategy
Anthropic has described its commitment as a unilateral move to have independent evaluators receive employee-like access. OpenAI would commit to embedding third-party evaluators and follow Anthropic’s same strategy.
How embedded oversight is supposed to work
The evaluators are intended to verify whether companies are following their pacing and safety commitments. The proposed arrangement would give them company badges, desks and laptops; in one reported case, OpenAI gave METR and Redwood roughly one week on its premises to investigate the Hugging Face incident. OpenAI also supports a provision in the FRONTIER Act that would allow independent verification organizations into top frontier labs.
The companies’ wider safety talks
OpenAI, Anthropic and Google DeepMind have been in talks about AI safety for several weeks. Amodei’s essay proposed a narrow government waiver allowing safety coordination, but the talks could risk violating antitrust law if the coordination were found to suppress competition.
The open-weights challenge
A competing policy argument calls for slowing AI progress while requiring any model offered to the public to be released as open weights. The proposed law would exempt internal and research models. Critics of that approach argue that funding for frontier-model development depends on valuations assuming the weights remain proprietary.
Why the pace debate persists
Amodei has warned that unchecked AI development may outrun the ability to understand and control these systems. He described AI’s rapid advance as being driven by its increasing ability to build the next generation of AI, or recursive self-improvement, and said he worries that such a swarm could be capable of taking over the entire internet within six to 12 months.