Business

Anthropic and OpenAI Propose Embedding Independent Safety Evaluators Inside AI Companies

2 min read

Anthropic CEO Dario Amodei proposed in a weekend essay that frontier AI companies embed third-party evaluators with the power to assess model alignment, report safety incidents, and publish findings without editorial control. Amodei committed Anthropic to giving evaluators such as METR and Redwood Research expanded access to its systems, and OpenAI CEO Sam Altman said his company would follow suit.

Evaluators broadly welcomed the proposal but said critical details remain unresolved, including how much access they will actually receive. Adam Gleave of FAR.AI said meaningful oversight would require access to intermediate training checkpoints, post-training environments, and employee interviews, rather than just finished models. Alexander Meinkeof Apollo Research said companies should be able to demonstrate whether models attempted to undermine their own alignment training, something currently unverifiable by outsiders.

Researchers noted past evaluations were often limited by time constraints, citing OpenAI’s week-long review of the Hugging Face incident and a three-day testing window for its GPT-6 Astra model. Henry Papadatos of Safer AI argued voluntary commitments remain fragile without binding regulation, since companies could reverse course during a public crisis.

Meta, Google DeepMind and SpaceXAI have not committed to the practice, though DeepMind’s Demis Hassabis has proposed a separate industry standards body. California’s SB 53 and SB 813, along with the EU AI Act, have begun establishing formal evaluation requirements.