Watching Agents Work: A Behavioral Audit of Offensive-Security LLM Runs
摘要
What closed and open models actually do when you tell them to hack a website Summary The cybersecurity capability of a model is currently measured by a solve rate, a percentage, and that number tells you almost nothing worth knowing. It doesn't tell you how the model behaved. It doesn't tell you what it is actually capable of, or where it is lacking, or how far it still is from doing the job end to end. And when you put an agent next to a human practitioner, the entire comparison collapses in
- 发布时间
- 收录时间
Skip to content