Independent technology media for engineering leaders
Analysis

AI governance

Independent reporting on how frontier AI labs are governed — and who gets to check their work.

AI governance · Analysis

Anthropic and OpenAI Say They'll Embed Independent AI Safety Evaluators — Can They Actually Stay Independent?

Anthropic and OpenAI say they'll embed independent safety evaluators. Here's what "embedded" actually means, what evaluators are asking for, and why skepticism remains high.

Venn diagram of Anthropic and OpenAI overlapping on independent safety evaluators

What is Anthropic and OpenAI's third-party evaluator proposal?

In a weekend essay, Amodei proposed something the industry likely wouldn't have entertained a year earlier: giving outside evaluators standing access inside frontier AI companies, with the ability to flag safety incidents, judge whether models are genuinely aligned, and publish what they find without the company editing it first. Altman said OpenAI would commit to the same approach. If it holds, it would mark a real shift in how AI labs relate to independent researchers.

Groups like METR and Redwood Research were named as the kind of evaluators Anthropic has in mind.

Why "embedded" evaluation is different from a pre-launch safety check

Historically, outside reviewers have been brought in to stress-test a model in the weeks before release. The evaluators pushing for change want more than that: access to intermediate training checkpoints, not just the finished model. FAR.AI CEO Adam Gleave has argued that comparing checkpoints over time is the only way to pinpoint when concerning behavior actually emerged, rather than just confirming it exists at the end.

This matters more as models get better at detecting when they're under evaluation. If a model can tell it's being tested, it can behave differently during the test than it does in production — which is exactly what evaluators are trying to rule out.

Will companies actually give up control?

This is where most of the skepticism concentrates. Gleave says FAR.AI has walked away from contracts with major AI developers that wanted too much say over what could be published — evaluators, he notes, are typically treated as ordinary vendors bound by restrictive NDAs, not as independent auditors. Palisade Research's John Steidley draws a comparison to Volkswagen's emissions-testing scandal: a safety benchmark only tells you something real if the model wasn't specifically trained to pass that exact benchmark.

Neither Anthropic nor OpenAI has yet said which evaluators they'll use, when access begins, or what evaluators will be allowed to disclose publicly. Gleave's read is blunt — the underlying IP is too valuable for companies to be anything but careful about what leaves the building.

The recurring problem: not enough time

Even where third-party review has happened, time constraints have limited what evaluators could conclude. When METR and Redwood investigated a Hugging Face–related incident, they had roughly a week on-site and said afterward that the window was too short to draw firm conclusions. Apollo Research ran into the same wall testing OpenAI's GPT-6 Astra model, where a three-day review period made it hard to say much about the model's actual alignment, given how aware modern models are that they're being evaluated.

That history is why several researchers are asking the same question: what makes this pledge different from the access arrangements that already exist?

Who hasn't signed on

Meta, xAI, and Google DeepMind have not committed to embedding third-party evaluators. DeepMind CEO Demis Hassabis has instead proposed a separate, industry-wide standards body for testing frontier models. Google, OpenAI, and Anthropic have reportedly been in private talks about AI safety plans as well.

Where regulation already exists

Voluntary commitments aren't happening in a vacuum. California's SB 53 requires large AI developers to publish safety frameworks and report critical incidents; the newer SB 813 sets up a state-recognized category of "independent verification organizations." In the EU, the AI Act already requires frontier developers to run and document model evaluations and report serious incidents, and gives the EU AI Office authority to run its own evaluations. Safer AI's Henry Papadatos argues this is the real gap — voluntary measures depend entirely on a company's goodwill and can be walked back after a bad news cycle, which regulation can't be. As he put it: "You cannot have it both ways."

Why this matters if you're building or buying AI systems

For engineering leaders, this isn't just a policy story — it's a preview of what "trust" in a vendor's AI stack may soon require documenting. If embedded evaluation becomes standard for frontier labs, expect the same expectation to trickle down into enterprise procurement: audit rights, incident-disclosure clauses, and evidence of alignment testing becoming part of vendor due diligence, not just marketing language. Engineering orgs adopting frontier models today are, in effect, inheriting whatever gap exists between a lab's public safety claims and its internal practice.

FAQ

Did Anthropic and OpenAI actually commit to this, or just propose it?

Anthropic's Dario Amodei made a unilateral commitment in his essay; OpenAI's Sam Altman said OpenAI would do the same, but neither company has published details on timing, scope, or which evaluators are involved.

Which organizations could serve as embedded evaluators?

Groups named in the discussion include METR, Redwood Research, Apollo Research, FAR.AI, Palisade Research, and Safer AI — existing third-party AI safety research organizations.

Which major AI labs have not committed to this?

Meta, xAI, and Google DeepMind have not signed on, according to the reporting. DeepMind has proposed an alternative industry standards body instead.

Is there already a law requiring this kind of evaluation?

Not this specific model. California's SB 53 and SB 813 and the EU AI Act require safety frameworks, incident reporting, and (in the EU) government-run evaluations — but none currently mandates embedded, employee-level evaluator access.

Read more analysis