Phying News
Curated security research, vulnerabilities, advisories and tools for practitioners.

How Anthropic’s Jailbreak Challenge Put AI Safety Defenses to the Test

Summary

Anthropic and HackerOne ran an AI red teaming challenge on a demo Claude 3.5 Sonnet to validate Constitutional Classifiers that block CBRN-related jailbreaks, awarding $55,000 to four teams.
Published
Collected

original ↗

Related coverage

back