Anthropic’s pledge to place third-party evaluators inside its operations is no longer presented as a solo move: OpenAI says it will adopt the same approach, shifting the proposal from a company policy to a shared test for frontier-AI oversight.
Anthropic’s oversight proposal becomes a two-company commitment
Anthropic has committed to using embedded third-party evaluators as part of its approach to pacing the AI frontier. The commitment was initially described as Anthropic’s own decision: the company was “unilaterally committing” to one of the strategies for pacing frontier development.
The proposed arrangement is not limited to an abstract promise of review. It would give evaluators company badges, desks and laptops, while OpenAI would adopt the same embedded-evaluator arrangement and follow Anthropic’s commitment to the strategy. Together, those points turn the proposal from a single company’s stated policy into a parallel commitment by two companies.
The idea of responsible development is being translated into physical access
Anthropic describes its mission as the responsible development and maintenance of advanced AI for the long-term benefit of humanity. Its evaluator pledge gives that broad mission a more concrete institutional form: the company has committed to embedded third-party evaluators rather than referring only to oversight as a general principle.
The practical proposal is notable because it places outside evaluators in the company’s working environment. Badges, desks and laptops describe access to the workplace and its tools, moving the idea of evaluation toward an embedded working arrangement. That makes the proposal operationally legible, even though the record does not establish what evaluators would be permitted to do once inside.
The record establishes participation, not yet the boundaries of the evaluators’ power
The admitted reports establish who is expected to participate and the basic workplace access proposed for them: third-party evaluators would be embedded, with company badges, desks and laptops, and OpenAI would follow the same model. They do not establish the evaluators’ authority, remit or decision-making power. Nothing in the record specifies whether evaluators could halt a development process, require changes, publish findings or otherwise act beyond conducting evaluations.
That boundary matters because physical access and institutional authority are different parts of an oversight system. The proposal supplies details about presence and equipment, while the parallel commitment supplies a second company adopting the arrangement; neither claim answers how the evaluators’ judgments would affect the companies’ conduct. The central unresolved issue is therefore not whether embedded participation has been proposed, but what power that participation would carry in practice.