ThinkFacility

News

The outside evaluators the labs promised to embed have had a standard ready since December 2025.

The AI Evaluator Forum welcomed the labs' new promises and pointed at five conditions it published before any of them were made.

On September 14, 2026, the people who would actually do the inspecting answered the AI labs. The AI Evaluator Forum posted at 12:10 pm Eastern, welcoming "recent statements regarding the importance of embedded independent experts to verify AI safety and security claims" and calling that "a first step towards trustworthy oversight of frontier AI".

Then it pointed at a document it had already written. AEF-1 is dated December 4, 2025, more than nine months before the pledges it's being offered to. The terms were on the table before the pledges.

Standard
AEF-1, Minimum Operating Conditions for Independent Third Party AI Evaluations
Version
Version 1, updated December 4, 2025
Covers
Access, conflicts of interest, autonomy, transparency, sensitive information
Still needed
Adoption by frontier developers

What the inspectors say they need

A long curved glass office building with a tall orange-paneled end section and a banner reading Next Generation EU, under a clear blue sky
The Berlaymont in Brussels, headquarters of the European Commission, whose AI Office has endorsed key provisions of AEF-1. Photo: Trougnouf (Benoit Brummer), CC BY 4.0, via Wikimedia Commons

The standard has five headings, each with specific conditions. The access conditions include safe harbor and computational resources, and the autonomy conditions include scoping, direct access and editorial control. Under transparency it lists conditions called "No contingent release" and "Redaction disclaimer".

Conflicts of interest have their own heading, which covers contingent compensation, organizational control, recusals and separate agreements, along with a conflict-of-interest policy and disclosure.

The success of AEF-1 requires adoption by a range of stakeholders, including frontier developers.

From AI Evaluator Forum

Who has already worked this way

The page lists evaluations run under those conditions. METR's Frontier Risk Report, in May 2026, assessed misalignment risk from AI agents deployed inside Anthropic, Google, Meta and OpenAI, with "more direct access to non-public information and more editorial independence than previous external evaluation engagements". (METR is the example Amodei's essay reaches for when it describes embedded evaluators.)

Transluce published one in August 2026 on how 77 model variants respond to users in a mental health crisis. SecureBio has been doing pre-release biosecurity assessments, most recently on GPT-5.6 Sol. The European AI Office has endorsed key provisions as a route for providers to meet the independence rules in the General-Purpose AI Code of Practice.

Where it rubs against the promise

Anthropic's commitment keeps a redaction right: security-sensitive, legally privileged, commercially sensitive or third-party confidential material can come out of a report before publication. Amodei's essay limits that right.

we can't redact findings just because they are unfavorable

From Dario Amodei — We Must Pace the Frontier

AEF-1 has its own conditions for redactions and for a redaction disclaimer, under the same transparency heading.

The page doesn't name any frontier developer as having adopted the standard. Its last line asks organizations to sign a public letter calling for greater transparency in third-party AI evaluations.