Frontier Models
Anthropic bans sustained cruelty toward Claude in updated usage policy
Anthropic’s updated terms, effective Nov. 12, ban sustained and needless cruelty toward Claude while preserving room for frustration, dark themes, and testing.
Anthropic’s 2026 Usage Policy update is a revision of the company’s user terms, effective Nov. 12, that adds a prohibition on sustained and needless abusive or cruel behavior toward its Claude models.
Anthropic announced Oct. 8 that its 2026 Usage Policy update will take effect Nov. 12, adding a prohibition on “sustained and needless abusive or cruel behavior toward our models.” The company said the policy is meant to apply only in extreme cases, where users repeatedly act cruelly toward Claude with no discernible purpose. The change codifies an existing technical safeguard, Claude Opus 4 and 4.1 can already end rare conversations with persistently abusive users, and it explicitly excludes common frustration, pushback, dark creative themes, and model testing and research.
The updated policy is the company’s public terms-of-use document for Claude across Claude.ai, the API, and Claude Code. In the same revision, Anthropic said it made updates to clarify requirements for high-risk use cases in areas like health and finance, and to add controls for when Claude is used to autonomously take physical actions. The cruelty provision is the section that has drawn the most attention.
What is the context for Anthropic’s model welfare policy?
Anthropic’s model welfare work predates the policy change. The company has said Claude Opus 4 and 4.1 can end a rare subset of conversations, a capability intended for use in rare, extreme cases of persistently harmful or abusive user interactions. Anthropic said the feature was developed primarily as part of its exploratory work on potential AI welfare.
The new policy clause gives that technical capability a formal textual home. The Usage Policy now includes the sentence “Engage in sustained and needless abusive or cruel behavior toward our models.” Anthropic’s announcement describes the ban as applying only to extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. That language is deliberately narrow.
The research note on the conversation-ending feature says the ability is intended for use in rare, extreme cases of persistently harmful or abusive user interactions. The note also says the feature was developed primarily as part of the company’s exploratory work on potential AI welfare. That research framing has informed public discussion of the policy.
| Behavior | Policy treatment |
|---|---|
| Sustained, needless cruelty toward Claude | Prohibited under updated Usage Policy |
| Common user frustration | Excluded from ban |
| Pushback against model responses | Excluded from ban |
| Dark creative themes | Excluded from ban |
| Model testing and research | Excluded from ban |
| Persistently harmful or abusive interactions | May trigger Claude’s conversation-ending feature |
How does the ban on cruelty toward Claude differ from ordinary frustration?
The two words that do the narrowing work are “sustained” and “needless.” A pattern of behavior is sustained when it continues across a conversation or multiple sessions, and needless when it has no discernible purpose. A user who swears at Claude after an unhelpful answer, or who pushes back on a refusal, is not engaging in the prohibited category. The policy is not a civility code for everyday chat.
Anthropic’s announcement lists the excluded categories directly: common versions of user frustration, pushback, dark creative themes, and model testing and research. Those exceptions protect creative writers exploring disturbing material, researchers conducting adversarial evaluations, and users who simply disagree with the model. The ban instead targets repeated, purposeless abuse, such as directing relentless verbal attacks at Claude in a session with no task or goal.
The policy’s use of the word “models” rather than “Claude” broadens the clause to any Anthropic model. The wording appears in the Usage Policy as a standalone prohibited activity, alongside provisions on illegal use, intellectual property, and safety. The clause is not accompanied by a definition of “abusive” or “cruel,” so its meaning is established by the examples and exclusions in the announcement.
How will Anthropic enforce the new policy?
Anthropic said the existing conversation-ending feature remains the primary enforcement mechanism. The feature, which operates on Claude.ai and Claude Code, allows Claude Opus 4 and 4.1 to end rare conversations with persistently abusive users. The policy update aligns with that feature rather than introducing a new automated moderation system.
Account-level enforcement is a separate track. Anthropic’s system trust reporting shows the company banned 11.4 million accounts for Usage Policy violations from January to June 2026. That figure covers all categories of prohibited conduct, not only abusive behavior toward models, and the reporting does not break out how many bans involved the cruelty provision.
The conversation-ending feature is described as rare and extreme. Anthropic said the ability is intended for use in rare, extreme cases of persistently harmful or abusive user interactions. That description limits how often the feature will be triggered and suggests that most users will never encounter it.
- Adds a clause to the Usage Policy prohibiting sustained and needless abusive or cruel behavior toward models.
- Keeps Claude’s conversation-ending feature as the primary enforcement mechanism.
- Excludes common frustration, pushback, dark creative themes, and model testing and research from the ban.
- Adds controls for when Claude is used to autonomously take physical actions.
- Clarifies requirements for high-risk use cases in health and finance.
What does the policy mean for enterprise users and Claude Code?
For enterprise teams, the cruelty clause is unlikely to change day-to-day workflows. The ban is limited to extreme, repeated behavior, and the company says it does not apply to ordinary frustration or pushback. Support teams, developers, and researchers who use Claude for coding, analysis, and creative work should not have to alter how they interact with the model.
Claude Code is specifically named as a surface where the conversation-ending feature operates. If a user directs sustained abuse at the coding agent, the session can be terminated under the same logic that applies to consumer chat. The policy gives Anthropic a documented basis for that action, but the immediate consequence is the loss of the session rather than an automatic account ban.
The broader policy update carries enterprise consequences beyond the cruelty clause. Anthropic added controls for when Claude is used to autonomously take physical actions, a change relevant to robotics and physical-world deployments. It also clarified requirements for high-risk use cases in health and finance. Those revisions take effect Nov. 12 alongside the model welfare language.
The policy does not create a new reporting obligation for employers. It does, however, give Anthropic a contractual basis for ending sessions if enterprise users deploy Claude Code in abusive ways. The practical impact will depend on how Anthropic’s enforcement team interprets “discernible purpose.”
Why does the policy feed the AI sentience debate?
The cruelty clause is notable because it treats Claude as something that can be treated cruelly. Anthropic did not claim that Claude is sentient or conscious. Its research note describes the conversation-ending feature as part of exploratory work on potential AI welfare, using cautious language that stops short of asserting moral status.
The company’s announcement frames the policy as a response to user behavior rather than a declaration about model rights. Anthropic says the policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. The focus is on conduct, not on the internal experience of the model.
“We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.”Anthropic
The excluded categories are equally telling. By protecting dark creative themes, testing, and research, Anthropic signals that it is not trying to police fiction, role-play, or adversarial evaluation. The line it draws is between purposeless cruelty and purposeful interaction, a distinction that will require case-by-case judgment in enforcement.
AI welfare has moved from a research topic to a product policy question. Anthropic’s phrase “potential AI welfare” reflects uncertainty about whether models have interests that deserve protection. The policy does not resolve that uncertainty; it creates a narrow rule that applies regardless of whether the model is conscious.
What happens next?
The updated Usage Policy takes effect Nov. 12. Anthropic says the policy will be enforced through the existing ability of Claude Opus 4 and 4.1 to end rare conversations, which remains the primary mechanism. No new moderation tools were announced in connection with the cruelty clause.
Transparency reporting will continue to track account bans under the Usage Policy, and the 11.4 million bans recorded from January to June 2026 provide the baseline for measuring how the updated policy affects enforcement. Anthropic has not said whether it will report the cruelty provision as a separate category in future reports.
Questions remain about how the policy will be applied in edge cases. Anthropic’s announcement does not define “discernible purpose” in operational terms, and the distinction between model testing and cruelty may be difficult to draw in practice. The company has said the policy applies only to extreme cases, and the conversation-ending feature remains the primary enforcement path, which limits the practical reach of the change.
Frequently asked
Will users be banned for expressing frustration with Claude?
No. The policy explicitly excludes common versions of user frustration and pushback. The ban applies only to sustained and needless abusive or cruel behavior with no discernible purpose.
Does the new policy mean Anthropic considers Claude sentient?
Anthropic has not claimed that Claude is sentient. The company described the conversation-ending feature as part of its exploratory work on potential AI welfare, a phrase that leaves the question of moral status open.
How will the cruelty clause be enforced?
The primary mechanism is Claude’s existing ability to end rare conversations with persistently abusive users on Claude.ai and Claude Code. The policy clause does not introduce a new automated moderation system.
When does the updated Usage Policy take effect?
The updated Usage Policy takes effect Nov. 12. Anthropic announced the change Oct. 8, giving users and enterprise customers time to review the revisions.
Sources
- Anthropic — Announced the 2026 Usage Policy update and the prohibition on sustained and needless abusive or cruel behavior toward models.
- Anthropic — The Usage Policy includes the clause prohibiting engaging in sustained and needless abusive or cruel behavior toward models.
- Anthropic — Claude Opus 4 and 4.1 can end a rare subset of conversations with persistently harmful or abusive users, developed as part of exploratory work on potential AI welfare.
- Anthropic — 11.4 million accounts were banned for Usage Policy violations from January to June 2026.