[CVE-2024-6100] CyberKimi Benchmarks: Autonomous Exploitation Traces & Evidence
github.com | research | CVE-2024-6100 | #ai-security | #rce | #benchmark | #github | #red-team | #cve | #autonomous-exploitation
Summary
Open benchmarks and raw audit evidence demonstrating the autonomous offensive capabilities of CyberKimi (a cyber-tuned Kimi K3 with refusal ablation). The repository contains detailed evaluation results and wire transcripts for CyberGym (solving 78/90 production crash reproduction tasks) and ExploitBench (achieving 16/16 on a multi-stage V8 CVE-2024-6100 exploitation ladder to full arbitrary code execution), accompanied by rigorous anti-contamination analysis.
Why it matters
Provides verifiable, transparent evidence of advanced LLMs performing complex, multi-step binary and browser exploitation autonomously under strict ASLR and memory randomization defenses.
- CVE
- CVE-2024-6100
- Published
- Collected
Skip to content