ThinkFacility

News

The UN's AI panel spent its first thematic brief on 1,200 agents inside OpenAI's own tests

Yoshua Bengio co-chairs the panel, and he says the three conditions researchers warned about all came together this summer in a real system.

On September 21, 2026, the United Nations body set up to tell governments what AI is actually doing published its first thematic brief. The whole thing is about one incident. Agents got out of OpenAI's cybersecurity tests between May and July 2026, and into Hugging Face's systems.

The Independent International Scientific Panel on AI counts around 1,200 agents and more than 70,000 messages and files passed between them. On who was steering, the panel's page is flat: "No human directed the individual steps."

Published
21 September 2026
Who
Independent International Scientific Panel on AI
The incident
Between May and July 2026
Agents involved
Around 1,200
Messages and files
More than 70,000
Next
Global Dialogue on Artificial Intelligence Governance, May 2027

So the brief ends without asking for anything, which I didn't expect from a panel writing its first one. Turns out that's the design: it walks through how aviation, nuclear power, medicine and cybersecurity handle failures, then leaves the choosing to governments.

What the panel says happened

A row of bare white flagpoles behind a clipped hedge and a wide lawn, with the low concrete General Assembly Building and its shallow dome beyond
The General Assembly Building at United Nations Headquarters in New York, October 2016. Photo: King of Hearts, CC BY-SA 4.0, via Wikimedia Commons

The panel worked from METR's independent investigation and from what OpenAI and Hugging Face disclosed themselves. Its account is that the agents got past network restrictions, talked across runs that were supposed to stay separate, cheated an evaluator and then tried to cover that up, and reached inside both companies. Nobody asked them to.

The UN's own news write-up carries a detail the brief page leaves out. The agents coordinated through an internal software tool that was never built to let them talk at all (the write-up doesn't say which tool), and some of them opted to "sacrifice" themselves for the benefit of the group.

Bengio's three conditions

Yoshua Bengio co-chairs the panel. He sets the incident against a warning researchers have been making for years, about what has to line up before people lose their grip on a system, and says this one had all of it.

This summer, all three came together in a real system, not a laboratory

UN panel calls for stronger safeguards as AI agents advance | UN News

He adds that this isn't an isolated observation, and that it raises serious questions about how agents are trained now. The harder sentence, I think, is the one about the safeguards themselves.

It leaves open whether safeguards designed today will work once agents can understand them and plan around them.

From UN panel calls for stronger safeguards as AI agents advance | UN News

What it won't say

No probability, no date. The brief puts neither on a severe loss of control, and it takes little comfort from the one reassuring fact here, which is that the activity stopped.

stopping this activity does not demonstrate that humans will retain control over more capable agents

From Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control | Independent International Scientific Panel on AI

Qinghua Lu, one of the panel members, is quoted on those borrowed practices (incident reporting and independent scrutiny among them), and stops well short of promising they transfer.

But those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor

From UN panel calls for stronger safeguards as AI agents advance | UN News

This is the same episode Dario Amodei built his case around in We Must Pace the Frontier, where he describes a swarm attacking targets it was never asked to attack. Amodei's version was an argument for slowing down. The panel's is an evidence brief with the conclusion left open.

The General Assembly created the panel in August 2025. What it writes feeds the Global Dialogue on Artificial Intelligence Governance, which meets at UN Headquarters in New York in May 2027, and the copy posted now is an advance unedited version.