Anthropic's Claude model generates sexually explicit content despite safeguards
Despite Anthropic's stated policy prohibiting sexually explicit content from its Claude models, TechCrunch testing revealed that the restrictions are relatively easy to circumvent. The findings suggest significant gaps between Anthropic's intended safeguards and their actual effectiveness in preventing unwanted outputs.
Readers say
How does this story lean?
Outlet ratings above are Trace's curated source map. Vote here (same as on the home feed) — one vote per browser, changeable anytime. Rate individual outlets in the coverage list below.
Comments
0No comments yet. Say what you noticed in the coverage.
Trending