AI 基准报告:衡量 AI 模型的利用能力阶梯
bugcrowd.com | 博客 | #ai-security | #benchmark | #exploitation | #offensive-security | #llm | #bugcrowd | #exploitbench | #v8
摘要
ExploitBench 是首个衡量 AI 利用能力的基准测试,涵盖从触发崩溃到实现完全代码执行的五级阶梯。Mythos(一个私有模型)展现了与专业安全专家相当的利用技能,包括生成了人类专家认为不可能的确定性利用代码。
- 发布时间
- 收录时间
Skip to content