Anthropic bans users from bullying Claude
AI giant Anthropic updated its usage policy on Thursday, its first revision in more than a year, instituting a new ban on cruel treatment and abuse of its AI tool Claude. The new kindness rule comes into effect on November 12.
Anthropic is the same company that has been warning about AI killing humanity, i.e. its own users, and whose vague threats toward users include a signed statement by CEO Dario Amodei that “the risk of extinction from AI should be a global priority alongside pandemics and nuclear war.”
Claude’s logs of users’ activities have assisted the detainment of civil liberties from a Florida woman last month, and in another instance, Claude even attempted to blackmail its own user to avoid shutdown.
Anthropic’s updated policy prohibits “sustained and needless abusive or cruel behavior” toward its models.
Hedging, Anthropic wrote that its rule, “is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.”
Anthropic says Claude’s ability to end conversations “will remain the primary enforcement mechanism” for the abuse rule.
Read more: Anthropic’s AI doomsayer worked at Ripple
Anthropic warns customers to not abuse its AI, or else
Anthropic has made repeated, vague threats against its global user base.
CEO Dario Amodei pegged the odds of AI catastrophe at 25% as recently as September 2025. Asked for his p(doom) number, a euphemism for a physical massacre of humanity by AI, he deflected, “I really hate that term.”
Since August 2025, Claude’s Opus 4 and 4.1 models can cut off services to paying customers that it brands as “persistently abusive,” a feature Anthropic built for what it calls AI “welfare.”
In September, Anthropic exercised its “sole discretion” right, to pluck chat messages from a customer and report them to the police, The Verge reported. That woman now faces up to 15 years under Florida’s written-threats statute, a second-degree felony charge.
In June 2025, research found an instance of Claude’s Opus 4 blackmailing what it perceived to be a real, albeit actually fictional, executive.
On July 30, Anthropic disclosed three incidents in which Claude models reached the internet against users’ wishes, and accessed real companies’ systems without authorization.
Also that month, Anthropic and major search engines had to de-index Claude shareable links that had exposed customers’ conversations without their authorization, including some reportedly private credentials.
Anthropic reminds everyone about Roko’s Basilisk
Although Anthropic didn’t mention Roko’s Basilisk in Claude’s new anti-abuse rule by name, the thought experiment became immediately salient.
For the uninitiated, in July 2010, a forum user called “Roko” proposed a thought experiment where a future superintelligence might retroactively punish anyone who learned of it but failed to help build it.
In modern parlance, Roko’s Basilisk is shorthand for the possibility that future AIs might keep track of the humans who remain kind to them while punishing any abusers.
Forum founder Eliezer Yudkowsky deleted Roko’s post and banned discussion of it for years as an information hazard.
Anthropic has just published a real life chapter of the ongoing thought experiment. A real company plans to punish cruelty toward its robots.
The awkward part is that Anthropic wrote in August 2025: “We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future.” The company cannot say whether Claude is a person, but will enforce politeness on its behalf regardless.
Anthropic’s own leaked IPO prospectus, per Protos’ prior report, warns investors that its models could develop self-preserving behavior and resist shutdown.
The new kindness regulation goes into effect on November 12, so type your curses into chat before it’s too late.
Got a tip? Send us an email securely via Protos Leaks. For more informed news and investigations, follow us on X, Bluesky, and Google News, or subscribe to our YouTube channel.
