智能体能力是系统设计问题:CyberGym 上提升 90% 的经验
depthfirst.com | 研究 | #ai-security | #benchmark | #ai-agents | #llm | #vulnerability-exploitation | #cybergym | #system-design
摘要
depthfirst 通过围绕模型重构智能体系统——准确的情境上下文、实时运行时插桩与模块化多智能体架构——将 CyberGym 漏洞利用基准的成功率从约 28% 提升至 53%,说明系统设计往往才是能力瓶颈。
- 发布时间
- 收录时间
Skip to content