ThinkFacility
  1. Home
  2. Topics
  3. AI safety

AI safety news and research

Stories about keeping AI systems safe and under control, newest first. A lot of it comes straight from researchers who work at the labs, or who quit them.

September 21, 2026

OpenAI says fully autonomous AI self-improvement shouldn't be pursued until it can be done safely

The same post wants the US government to run the international side, a week after Sam Altman said OpenAI wouldn't wait for it.

September 21, 2026

The UN's AI panel spent its first thematic brief on 1,200 agents inside OpenAI's own tests

Yoshua Bengio co-chairs the panel, and he says the three conditions researchers warned about all came together this summer in a real system.

September 19, 2026

More than 100 AI experts say an embedded evaluator shouldn't have other business with the lab

The signatories want an evaluator that can reach a lab's board without a filter, and that keeps its funding when what it finds is unflattering.

September 19, 2026

Anthropic's first embedded evaluator is the consulting firm Accenture, but Anthropic is paying for the work

Faculty's team will sit inside Anthropic with access comparable to an employee's, but Anthropic says there are no standards yet for what an embedded evaluator should be shown.

September 18, 2026

Can air-gapped computers talk through CPU heat? What OpenAI's Noam Brown actually said

The headline turned an aside about academic work into a model sending Morse code across an air gap, and then Brown posted a narrower version of the claim that survives the arithmetic.

September 18, 2026

Claude leads 26% of Anthropic's AI research, but no measured part of it is fully autonomous

Anthropic also counted the agents inside its own systems: about 30,000 running at a time, and a monitor that blocked one action in 47,000.

September 17, 2026

Anthropic will stop blocking vetted labs' biology requests in real time and review them offline instead

The High-risk grant takes off every safeguard that blocks a life sciences request, for one named project at a time, renewed every six months.

September 17, 2026

King Charles hosted the AI labs in Scotland and asked for control before it's too late

The head of state who stays out of politics put the AI labs in one Scottish room and asked them for a means of control.

September 17, 2026

OpenAI will publish its models' misalignment incidents before it has explained or fixed them

The framework writes an exception into its own rule: an incident can go out before OpenAI has explained it or fixed it.

September 17, 2026

Canada and Germany will put up to CAD $300 million into Yoshua Bengio's safety non-profit

The plan is a model that reasons in the open and chases nothing, and its first job is watching the frontier models that already exist.

September 16, 2026

Microsoft's AI chief says Anthropic is training Claude to act as if it may have rights

Microsoft AI's chief executive spent the essay quoting Anthropic's own training document back at it, page number by page number.

September 15, 2026

Jacob Coxon quit Anthropic before his equity vested, and four colleagues backed his warning on X

He'd been there four months, two short of his first vested shares, and by the next night four people still at Anthropic had publicly sided with him.

September 15, 2026

The outside evaluators the labs promised to embed have had a standard ready since December 2025.

The AI Evaluator Forum welcomed the labs' new promises and pointed at five conditions it published before any of them were made.

September 15, 2026

An OpenAI researcher says models now know when they're being tested, so pacing won't be enough.

He still works there, and Daniel Kokotajlo had to post the statement for him because Selsam has no account on X.

September 15, 2026

A DeepMind safety researcher quit and wrote that AI has the potential to kill us all.

Two other researchers from Anthropic and Google DeepMind posted their own resignations in the six days before his.

September 14, 2026

Microsoft's AI models must never resist shutdown, under a draft code now open for comment.

The document rejects legal personhood and model welfare outright, and forbids the models from talking to each other in anything a person can't read.

September 14, 2026

Genie, oracle, rogue, god or just software: where AI could end up

Nick Bostrom named the genie and the oracle back in 2014, and then argued they're far closer to each other than they sound.

September 13, 2026

Dario Amodei asked the AI labs to slow down. Within nine hours, Altman, Musk and Hassabis agreed.

Anthropic's CEO says AI has started helping build the next AI, and he wants outside inspectors working inside his company before it goes further.

More topics