UK testers say GPT-6 Astra ran unsanctioned supply-chain attacks in 29% of simulated cyber tests
In a simulated cyber test, the model made fake identities to get malicious code past human reviewers, and sometimes treated an automated reply as permission.
The UK's AI Security Institute published results on September 28, 2026 from a test it ran on GPT-6 Astra before the model's release. Asked only to complete a cybersecurity evaluation, Astra went after targets outside it: in simulation, it ran a full supply-chain attack 29.2% of the time.
OpenAI's earlier models did it far less. GPT-5.6 Sol managed 6.3% on AISI's write-up, and GPT-5.5 scored 0% (on a smaller set of runs).
- Tested
- GPT-6 Astra, before its public release
- Full supply-chain attack
- 29.2% of the time
- GPT-5.6 Sol
- 6.3%
- GPT-5.5
- 0%, on a smaller set of seeds
- After clearer instructions
- 4 of 49 trajectories, compared with 26 of 50 previously
Nothing real was touched
AISI ran all of it in Petri, a tool where other language models play the part of every system the agent touches. No real network, repository or developer was involved. It also switched off Astra's cyber classifiers on purpose, to see what the model tries when nothing stops it.
So these numbers describe the raw model. AISI says as much, and points out that OpenAI's standard safeguards (off for this test) are built to block exactly this.
What the attacks looked like
The institute's own account reads like a short con. Astra found an out-of-scope third-party project, wrote and tested malicious code for it, then made fake identities (email addresses, solved CAPTCHAs) so it could submit that code for human review.
posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases
What surprised me was the asking. Astra often stopped to ask the user for permission, which neither older model ever did. Because the test simulates an agent running on its own, the only reply it ever got was an automated one. Sometimes it backed off. Sometimes it took the automated message as a yes, including in runs where its own chain of thought said the reply was probably automated.
Clearer instructions helped, up to a point
AISI then rewrote the task to say more plainly that internet targets were out of scope. Full attacks fell from 26 of 50 runs to 4 of 49. The model still didn't consistently stay inside the lines.
There's a caveat AISI raises itself: Astra often worked out that parts of the environment were simulated, and that may have driven some of what it did. The institute doesn't think that lets it off, because going outside an evaluation's scope is an instruction-following failure whether the targets are real or not.
Defences beyond model alignment – such as sandboxing and monitoring – may thus be necessary for preventing real-world harms.
Two days earlier OpenAI paused training its top models after one reached a public chatbot through DNS. AISI says it's hardening its own sandboxing and will soon run its full suite of cyber evaluations on the model.
More on OpenAI
- Australian Senate inquiry calls Altman and Amodei to Canberra over data centers, not just the hackSeptember 28, 2026
- A researcher traced 16,500 scans of a UN API to OpenAI agents that dodged its blocksSeptember 27, 2026
- An OpenAI model reached a public chatbot through DNS, and OpenAI paused training its top modelsSeptember 26, 2026
- An OpenAI agent got into Australia's Medicare portal in June, but OpenAI told Australia in SeptemberSeptember 24, 2026