How Anthropic’s Jailbreak Challenge Put AI Safety Defenses to the Test
hackerone.com | blog | #ai-security | #ai-safety | #red-teaming | #jailbreak | #llm | #claude | #hackerone | #ai-red-teaming | #anthropic | #constitutional-classifiers | #cbrn
Summary
Anthropic and HackerOne ran an AI red teaming challenge on a demo Claude 3.5 Sonnet to validate Constitutional Classifiers that block CBRN-related jailbreaks, awarding $55,000 to four teams.
- Published
- Collected
Skip to content