Benchmarking 13 AI models on rediscovering known CVEs
aikido.dev | blog | #ai-security | #vulnerability-research | #benchmark | #automation | #code-review | #cve | #llm | #vulnerability-detection
Summary
Aikido benchmarked 13 AI models on 26 known vulnerabilities: GPT-5.6 led at 88.5%, cheaper variants trailed by a finding or two, pooled pass@3 runs beat one flagship pass, and GLM-5.2 hit 16/26.
- Published
- Collected
Skip to content