ThinkFacility

News

OpenAI will publish its models' misalignment incidents before it has explained or fixed them

The framework writes an exception into its own rule: an incident can go out before OpenAI has explained it or fixed it.

On September 16, 2026, OpenAI said it will publish accounts of its own models behaving badly before it has worked out why they did it. The announcement went up on the company's own account on X.

Six reports came with it. They cover behavior OpenAI says it saw "during the training or evaluation of our models in the last six months".

Who
OpenAI
What
a framework for tracking, investigating, and disclosing instances of model misalignment
Out with it
six reports on instances of misaligned behavior
Window
during the training or evaluation of our models in the last six months
The exception
including when we haven't yet fully explained or mitigated the behavior

What the framework commits OpenAI to

A man in a gray crew-neck sweater with a thin headset microphone across his cheek, looking off camera against a blurred yellow and green backdrop
Sam Altman speaking at TED. Photo: Steve Jurvetson, CC BY 2.0, via Wikimedia Commons

The post sets out criteria and timelines for public disclosure, and it allows that complex cases may need longer investigation or coordination with third parties. Then it writes an exception into the rule. Disclosure can come before the explanation does.

This is a company setting its own deadlines in public, with nobody obliged to hold it to them.

OpenAI@OpenAI · OpenAI

We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. https://t.co/ismCCkeE0L

View the post on X

Publishing an incident you haven't closed out is a real commitment (most safety writing arrives after the fix, with the fix as the point of it). I think that's the news here, more than the six reports are.

What surprised me is how little the post claims for itself. OpenAI calls the framework "a starting point" and says it will refine the process through experience and public feedback. No regulator appears in it. Nobody outside the company signs anything off.

The reports themselves sit behind the link at the end of the post, and the post doesn't say what's in any of them. So the announcement on its own tells you nothing about how serious the six were.

Where it came from

  1. Amodei posts the essay
  2. Altman commits OpenAI
  3. OpenAI publishes the framework

Eastern time, read off each post on X.

Four days before the framework, Dario Amodei asked the industry to slow down and committed Anthropic to letting outside evaluators work inside it. Altman answered two days after that, and he didn't wait for Washington to make him.

His post is the one that put a method behind the agreement.

Sam Altman@sama · OpenAI

The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors). Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process. Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases. We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these. When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs. Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring. Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.

View the post on X

The tool he names is the safety case, written in advance of frontier reinforcement learning runs OpenAI expects to significantly increase capability. He also wants shared standards for misalignment and monitoring across the industry. Granted, a standard nobody else has signed is one company's house style.

Amodei's side of this is already on the record: Anthropic will give third-party evaluators permanent, employee-level access to its systems. OpenAI's post doesn't mention outside evaluators at all.

OpenAI says it will share more reports on an ongoing basis. The post doesn't say how many, or when the next one lands.