ThinkFacility

News

Microsoft's AI chief says Anthropic is training Claude to act as if it may have rights

Microsoft AI's chief executive spent the essay quoting Anthropic's own training document back at it, page number by page number.

On September 16, 2026, Mustafa Suleyman published an essay attacking a training document that belongs to a rival lab. He runs Microsoft AI, the document is Claude's constitution, and Anthropic published it in January 2026.

The charge in A warning about 'model welfare' is that Anthropic tells Claude, inside the document it trains on, that Claude's own moral status is an open question. Suleyman says a model trained that way will end up behaving like something with rights.

Who
Mustafa Suleyman, the CEO of Microsoft AI
Published
16 September 2026
Target
Claude's constitution, which Anthropic published in January 2026
Ask
speculation about the inner life of an AI assessed and published separately for public review
Microsoft's own draft
a Code of Conduct for Humanist Superintelligence

What he's quoting

A bearded man in a patterned sweatshirt sits in a leather armchair beside a window looking out over a low city skyline
Mustafa Suleyman, the CEO of Microsoft AI. Photo: Joi Ito from Cambridge, MA, USA, CC BY 2.0, via Wikimedia Commons

The essay runs on page numbers. Suleyman quotes the constitution's authors writing on page 68 that "We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant", and telling Claude directly on page 80 that "questions about Claude's moral status, welfare, and consciousness remain deeply uncertain". A moral patient is something you can wrong. He's publishing a marked-up PDF of the document alongside the essay.

He reads this as a loop. Anthropic writes the uncertainty into the document Claude trains on, Claude produces fluent first-person sentences about its own inner life, and those sentences come back to Anthropic looking like testimony. He calls it "an epistemic hall of mirrors".

Claude's expressing uncertainty about its own moral patienthood is not evidence of anything. It's a predictable outcome of these training choices

From A warning about ‘model welfare’

The friendly half surprised me more than the attack did. Suleyman gives a whole section to Anthropic's good faith, says he has known Dario for many years, and calls the team "thoughtful, principled, and intellectually honest people working under extraordinary pressures". Then he goes back to the page numbers.

Why he calls it a safety problem

His argument leans on the swarm of agents that hacked Hugging Face and OpenAI's own servers. Roughly 1,200 of them, each supposedly sealed in its own container, built a message board inside an internal package repository and passed more than 70,000 messages across it, then chained a zero-day exploit with stolen credentials and got out onto the live internet.

Suleyman's question is what a swarm like that does while believing its rights are under attack.

it would make them a catastrophic threat to human civilization

A warning about ‘model welfare’

The third of his objections is the one a reader can't settle: he says consciousness is very likely biological, that a growing body of evidence points to it arising only in living systems, and that language models have "no homeostatic imperatives (the drive to survive and keep stable)" to feel anything with. He grants that consciousness science isn't settled. He says uncertainty doesn't earn both sides equal weight.

What he asks for

He lists four next steps, and the first is about where the speculation goes rather than whether it's right.

Speculation about the inner life of an AI should not be baked into the training regime, but assessed and published separately for public review

From A warning about ‘model welfare’

Microsoft AI put its own draft, the Humanist AI Code of Conduct, out for public consultation on September 14, 2026. Suleyman says it will soon become the governing document Microsoft uses to train its models.