13 个 AI 模型在已知 CVE 重现测试中的基准评测
aikido.dev | 博客 | #ai-security | #vulnerability-research | #benchmark | #automation | #code-review | #cve | #llm | #vulnerability-detection
摘要
Aikido 用 26 个已知漏洞评测 13 个 AI 模型:GPT-5.6 以 88.5%(23/26)居首,廉价变体与旗舰差距仅一两处,多次运行汇总(pass@3)可靠超过单次旗舰;开源权重 GLM-5.2 达到 16/26。
- 发布时间
- 收录时间
Skip to content