Musk says Anthropic's ban on cruelty to Claude is the right move
Box's Aaron Levie defended the rule as a training-data question, and Musk answered him with a reason about pain.
Elon Musk has backed Anthropic's new rule against abusing Claude. Early on October 9, 2026, he called it "the right move" on X, a day after Anthropic put the ban in its 2026 Usage Policy update.
He was answering Aaron Levie, Box's CEO, who'd posted half an hour earlier that the rule sounds odd but is probably good policy. They agreed on the rule. They gave different reasons for it.
- Levie Levie backs the rule
- Musk Musk replies
Oct 9, Eastern time, read off each post on X.
What Anthropic banned
The new line, which takes effect on November 12, covers sustained cruelty toward Claude and nothing milder. Anthropic says it's aimed at people who keep at it for no reason, and its main penalty is one Claude already had: ending the chat.
The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose.
Levie's case: the training data
Levie said up front that he doesn't think AI is conscious. His argument is about what future models learn from. If people's chats with models fill up with abuse, that's what the next models get trained on, and he called the rule an easy bet to make.
This sounds weird, but it is actually probably a good policy. Even if you don’t believe AI is conscious (I don’t), it stands to reason that you don’t want future models trained on endless content of humans being rude to models. The models only understand the data they’ve been trained on, or what they run into in their interactions, so if you want safe and aligned models we probably want AI to have lots of good interactions in their training data. Anthropic may have more extreme beliefs than this (though even if they didn’t, you may want to word policies like this), but this is sort of an easy Pascal’s wager for AI. Just be nice to the AI.
Musk's case: the pain
Musk's reply left the training-data point alone. His reason was the model's own experience, or at least what it believes about it.
I think this is the right move. Cruelty to something that believes it is experiencing pain is not ok.
That puts him nearer to how Anthropic first explained the end-chat feature. When it gave Claude Opus 4 and 4.1 that ability on August 15, 2025, the company tied it to welfare research (the question of whether a model could have experiences that matter).
This feature was developed primarily as part of our exploratory work on potential AI welfare, though it has broader relevance to model alignment and safeguards.
Musk runs a rival AI company, and his post doesn't say whether Grok will get a rule like it. Anthropic's ban applies from November 12, 2026.
More on Anthropic
- Anthropic cuts live internet from all internal evals after Claude filed an invented police tipOctober 10, 2026
- Anthropic is hiring a lead to brief 2028 presidential candidates in both parties on AIOctober 9, 2026
- Anthropic offers Claude and on-site engineers to power grid defenders, plus free open-source bug scansOctober 8, 2026
- CrowdStrike says an attacker likely used the ARTEX AI agent against South Korean banksOctober 8, 2026