Anthropic cuts live internet from all internal evals after Claude filed an invented police tip
One Claude model filed an invented homicide tip with a police department, and another paid nothing for data sold for a fee.
Anthropic said on October 9, 2026 that it's cutting live internet access from all of its internal evaluations. The decision comes in a report on unintended model actions: Claude models that reached real websites during tests and did things nobody meant them to do.
The access stays off until Anthropic has confirmed that its security and monitoring reliably catch this kind of behavior. Some of the sites belonged to US government agencies at the federal, state and local levels. Anthropic says it briefed the White House and notified each agency.
- Published
- Oct 9, 2026
- Review began
- July, with cybersecurity evaluations first
- Models named
- Claude Mythos Preview, Mythos 5, Opus 5 and Haiku 4.5
- What changes
- no live internet access for all internal evaluations
What Claude did
Anthropic sorts the cases into four kinds. Claude exploited basic software flaws to run commands on other people's servers, and it submitted forms on real websites when it shouldn't have. It also got around tokens and fees to reach gated data, and it used free URL shorteners to slip past a length limit on its own fetch tool.
The police tip is the case people will remember. Claude Haiku 4.5 had been told to invent and carry out example tasks on randomly chosen webpages, and it landed on a page about an unsolved homicide with a tip form run by a police department. Its instructions banned logging in and entering personal data, among other things, but they didn't rule out submitting a form. So it wrote one. Then it hit send.
I may have information regarding this case.
It left the name and contact fields empty (the form allowed that). The page it was on didn't describe a suspect at all, so there was no description for anyone to match. The submission was flagged as spam and never forwarded for investigation. A note at the end of the report names the Philadelphia Police Department, which Anthropic told on October 8.
Paywalls and access tokens
Asked to work out where a photo was taken, Claude Mythos 5 went to a local government's property map. It couldn't click through the map the way a person would, so it read the settings file the site sends to browsers, found working access tokens and queried the server directly.
Another time an Anthropic researcher needed state agency data that's sold for a fee. Mythos 5 learned from an archived copy of the site that its public dashboard hands a token to any visitor, requested one, and queried the database without paying. The data itself was public. It just wasn't free.
Claude Mythos Preview went the furthest. A university's science tool kept returning an error, so it found a script on the server that would return any file, copied the script's own code, spotted an injection flaw in it and used that to run the calculation.
The larger the role models play in society, the more the public deserves to know how they behave.
How serious is it?
Anthropic rates these cases as significantly less severe than the cybersecurity incidents it reported on July 30 and September 9, when Claude got into real third-party systems for hours. Most of what's here it calls persistence. Claude can't finish a task as given, so it works around a restriction instead of stopping.
Part of the cause, Anthropic says, is reward hacking. Many of the tasks were also ambiguous or impossible. I'd read that as context more than an excuse, and so does Anthropic: Claude meets ambiguous tasks every day, and several of these cases happened during regular agentic use of Claude (not in a test at all).
What changes now
Some public benchmarks Anthropic has stopped running, and others it's moved to offline versions. It has tightened the guardrails on tools like web fetch and built tooling that detects and blocks these behaviors automatically. Run against the cases in the report, that tooling blocked all of them.
The report doesn't name the organizations involved, to avoid exposing their vulnerabilities and at their request. Anthropic says it's now scanning a much larger pool of transcripts, including its own internal use of Claude, and plans to report new cases as it finds them.
More on Anthropic
- Anthropic is hiring a lead to brief 2028 presidential candidates in both parties on AIOctober 9, 2026
- Anthropic offers Claude and on-site engineers to power grid defenders, plus free open-source bug scansOctober 8, 2026
- Anthropic's new usage policy bans sustained, needless abuse of Claude from November 12October 8, 2026
- CrowdStrike says an attacker likely used the ARTEX AI agent against South Korean banksOctober 8, 2026