Back to news

OpenAI joins push for embedded independent evaluators after Anthropic safety call

Sam Altman committed OpenAI to giving external evaluators employee-like access, turning an industry warning about fast-moving frontier systems into a concrete governance proposal.

A rival company adopts the evaluation proposal

OpenAI has committed to giving independent evaluators employee-like access to its systems, joining a safety proposal advanced by Anthropic chief executive Dario Amodei. The move is a new corporate action following Amodei’s broader call to slow the unchecked advance of frontier artificial intelligence. Anthropic separately said it would provide permanent, employee-level access to third-party evaluators. Sky News reported the commitments on September 13.

The distinction between endorsement and commitment matters. Sam Altman said OpenAI would adopt the evaluator-access proposal and provide more detail later. Elon Musk supported the general warning and suggested peer review between competitors, but Sky reported that he stopped short of matching Anthropic’s commitment. No common timetable, evaluator-selection process or enforcement mechanism has yet been announced.

What embedded access could change

External evaluations usually test a model at a particular point, often through constrained interfaces. Employee-like access could let specialists observe systems more continuously, examine pre-release versions and compare safety claims with internal practices. In principle, that could improve incident detection and make voluntary commitments more auditable. Its value will depend on whether evaluators can publish adverse findings, obtain relevant technical information and operate without commercial pressure from the companies they assess.

Official UK research supports the need for rigorous testing while warning against dramatic extrapolation. The government-backed AI Security Institute found rapid improvement in several controlled capability tests and vulnerabilities in every system it examined. It also stressed that laboratory results are not predictions of real-world harm. That combination—measurable capability growth, persistent safeguards gaps and substantial uncertainty—helps explain the demand for continuing access instead of one-off certification.

Voluntary coordination still has limits

Amodei’s proposal extends beyond embedded evaluators. It calls for common standards among companies in democratic countries and eventual coordination with governments including China. Those are diplomatic and regulatory ambitions, not completed agreements. Even the narrower corporate pledge leaves unresolved questions about who qualifies as independent, how sensitive findings will be handled and whether access survives disagreements over product launches.

The next meaningful development will be implementation. OpenAI and Anthropic would need to identify evaluators, define their authority, disclose reporting arrangements and explain how findings can delay or alter deployment. Regulators will also decide whether voluntary access is sufficient or should become a legal requirement. Until those details appear, the commitments are a governance experiment rather than proof that the competitive race between frontier laboratories has slowed.