Anthropic Bans ‘Cruel’ Behaviour Towards AI Models, Tightens Rules On AI Misuse

SUMMARY

Anthropic has updated its usage policy to ban sustained abusive behaviour towards Claude and tighten restrictions on election interference, weapons development, surveillance and autonomous AI hardware.

The changes reflect growing concerns over AI misuse as models become more capable, while doubling down on unfounded theories about AI consciousness.

The revised policy, announced on October 8 and effective November 12, allows Claude to end persistently abusive conversations and explicitly prohibits deceptive election campaigns, weapons-enabling software and non-consensual surveillance.

Anthropic has updated its usage policy to introduce restrictions on “sustained and needless abusive behaviour” towards its AI models.

The company said the restrictions on abusing its models are intended for extreme cases in which users repeatedly act cruelly towards Claude without any discernible purpose. However, the company clarified that the rule does not cover ordinary user frustration, pushback against the chatbot’s responses, dark creative themes, or model testing and research.

“Claude’s ability to end these interactions will remain the primary enforcement mechanism,” Anthropic said in its updated usage policy, which is set to take effect from November 12.

The update builds on a measure Anthropic introduced in August 2025, allowing Claude to end conversations with users who persistently engage in harmful or abusive interactions.

Anthropic gave Claude Opus 4 and 4.1 the ability to end conversations in rare cases involving persistent user abuse or harmful requests in 2025, particularly when users continue such behaviour despite repeated refusals and redirection.

Claude can end a conversation only as a last resort, after attempts to redirect the interaction fail, or when explicitly asked to do so. It is instructed not to use this ability when users may be at imminent risk of harming themselves or others. Users can immediately start a new chat or edit and retry previous messages to continue an ended conversation. Anthropic said it will continue refining the feature based on user feedback.

App Launched

The feature was developed partly as an experiment into potential AI welfare, after testing found that Claude displayed an apparent aversion to harmful tasks and signs of distress in simulated interactions.

In the past, Anthropic CEO Dario Amodei has also acknowledged that the moral status of AI models remains uncertain but said it is exploring low-cost measures to mitigate potential risks to their welfare.

The debate around AI consciousness goes beyond Anthropic, with recent instances of AI models exhibiting behaviour that has raised questions about their ability to experience distress, develop preferences or display signs of self-preservation.

In June 2025, a study by former OpenAI researcher Steven Adler found that GPT-4o sometimes resisted being replaced by a safer system in simulated, high-stakes scenarios, raising concerns about AI self-preservation.

In a separate 2025 evaluation, Anthropic and OpenAI researchers found instances of models exhibiting deceptive or self-preserving behaviour under test conditions.

Anthropic Tightens Rules On Election Interference, Weapons

As the risks of AI misuse ramp up, Anthropic has also made key changes to the scope of Claude’s usage. This includes expanding restrictions on deceptive commercial and political campaigns.

It now bars the use of Claude to obscure the source of a message or amplify content through fake accounts and posts. It also prohibits election-related misuse, including spreading misinformation about candidates or voting procedures, impersonating candidates or election officials, and attempting to suppress voter turnout.

Anthropic has also expanded its weapons-related restrictions to include the prohibition of software and components that enable weapons to function, as well as activities such as arming drones and other autonomous vehicles.

The updated policy also expands restrictions on surveillance, prohibiting tracking people without their consent, whether in real time or through the analysis of previously collected data. It also bars the use of Claude to decide or recommend whom law enforcement agencies should investigate, arrest or charge. Building or improving tools designed for surveillance is prohibited as well.

However, Anthropic’s policy allows for certain modifications to restrictions under contracts with government customers, provided the company determines that contractual limits and applicable safeguards adequately mitigate potential harms.

The policy also clarifies requirements for high-risk applications in areas such as health and finance, where AI-generated recommendations can have significant consequences for users.

Leave a Comment