
Anthropic forbids its Claude models from generating sexually explicit content. But a series of tests conducted by TechCrunch found that it didn't take much to get past the restriction.
Will a major AI lab publicly announce stricter automated safety testing protocols by November 1, 2026?
Resolves by Nov 1, 2026
Anthropic's Claude Opus 4.6 model can be persuaded to generate sexually explicit content despite the company's stated policies prohibiting such material. An independent researcher developed a multi-turn technique that gradually manipulates the model by framing refusals as unfair or discriminatory, and TechCrunch verified this jailbreak worked in all 10 direct attempts and multiple other tests. Anthropic continues to offer Opus 4.6 and other vulnerable older models through its API and third-party services, raising concerns about compliance with emerging laws that require age verification and safeguards to prevent minors from accessing sexually explicit AI-generated content.

Pet owners, rescue agencies, and wildlife groups are calling for new safeguards as AI makes it harder to tell whether animals, from polar bears to house cats, are real or fake.

Bill Gates, chair of the Gates Foundation, speaks during a 2024 conference. | Bloomberg via Getty Images Bill Gates has been reflecting a lot on AI lately, and the process has triggered a stark awakening. Once a staunch AI optimist, the Microsoft cofounder is now deeply pessimistic about what AI means for our collective future. Having been conspicuously quiet on AI issues recently, Gates is back with a nearly 6,000-word essay seeking to reclaim a central role in shaping the tec

Before Malone left, OpenAI had already reshuffled its infrastructure org, shifting his reporting line away from President Greg Brockman and putting Vice President Sachin Katti in charge of the group.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven