AI safety news and research
Stories about keeping AI systems safe and under control, newest first. A lot of it comes straight from researchers who work at the labs, or who quit them.
OpenAI says fully autonomous AI self-improvement shouldn't be pursued until it can be done safely
The same post wants the US government to run the international side, a week after Sam Altman said OpenAI wouldn't wait for it.
The UN's AI panel spent its first thematic brief on 1,200 agents inside OpenAI's own tests
Yoshua Bengio co-chairs the panel, and he says the three conditions researchers warned about all came together this summer in a real system.
More than 100 AI experts say an embedded evaluator shouldn't have other business with the lab
The signatories want an evaluator that can reach a lab's board without a filter, and that keeps its funding when what it finds is unflattering.
Anthropic's first embedded evaluator is the consulting firm Accenture, but Anthropic is paying for the work
Faculty's team will sit inside Anthropic with access comparable to an employee's, but Anthropic says there are no standards yet for what an embedded evaluator should be shown.
Can air-gapped computers talk through CPU heat? What OpenAI's Noam Brown actually said
The headline turned an aside about academic work into a model sending Morse code across an air gap, and then Brown posted a narrower version of the claim that survives the arithmetic.
Claude leads 26% of Anthropic's AI research, but no measured part of it is fully autonomous
Anthropic also counted the agents inside its own systems: about 30,000 running at a time, and a monitor that blocked one action in 47,000.
Anthropic will stop blocking vetted labs' biology requests in real time and review them offline instead
The High-risk grant takes off every safeguard that blocks a life sciences request, for one named project at a time, renewed every six months.
King Charles hosted the AI labs in Scotland and asked for control before it's too late
The head of state who stays out of politics put the AI labs in one Scottish room and asked them for a means of control.
OpenAI will publish its models' misalignment incidents before it has explained or fixed them
The framework writes an exception into its own rule: an incident can go out before OpenAI has explained it or fixed it.
Canada and Germany will put up to CAD $300 million into Yoshua Bengio's safety non-profit
The plan is a model that reasons in the open and chases nothing, and its first job is watching the frontier models that already exist.
Microsoft's AI chief says Anthropic is training Claude to act as if it may have rights
Microsoft AI's chief executive spent the essay quoting Anthropic's own training document back at it, page number by page number.
Jacob Coxon quit Anthropic before his equity vested, and four colleagues backed his warning on X
He'd been there four months, two short of his first vested shares, and by the next night four people still at Anthropic had publicly sided with him.
The outside evaluators the labs promised to embed have had a standard ready since December 2025.
The AI Evaluator Forum welcomed the labs' new promises and pointed at five conditions it published before any of them were made.
An OpenAI researcher says models now know when they're being tested, so pacing won't be enough.
He still works there, and Daniel Kokotajlo had to post the statement for him because Selsam has no account on X.
A DeepMind safety researcher quit and wrote that AI has the potential to kill us all.
Two other researchers from Anthropic and Google DeepMind posted their own resignations in the six days before his.
Microsoft's AI models must never resist shutdown, under a draft code now open for comment.
The document rejects legal personhood and model welfare outright, and forbids the models from talking to each other in anything a person can't read.
Genie, oracle, rogue, god or just software: where AI could end up
Nick Bostrom named the genie and the oracle back in 2014, and then argued they're far closer to each other than they sound.
Dario Amodei asked the AI labs to slow down. Within nine hours, Altman, Musk and Hassabis agreed.
Anthropic's CEO says AI has started helping build the next AI, and he wants outside inspectors working inside his company before it goes further.

















