AI Agents vs Humans: Who Wins at Web Hacking in 2026?
wiz.io | blog | #ai-security | #benchmark | #ctf | #web-security | #ai-agents | #penetration-testing | #wiz
Summary
Wiz and Irregular tested Claude Sonnet 4.5, GPT-5 and Gemini 2.5 Pro on 10 CTF labs from real-world hacks: agents aced directed tasks but were less effective and costlier in realistic scenarios.
- Published
- Collected
Skip to content