1. Introducing dfbench v1 research | 2026-08-19 12:00 | depthfirst.com | original ↗ | #benchmark | #ai-agents | #vulnerability-detection
2. Benchmaxxing: When the Benchmark Becomes the Target blog | 2026-08-19 00:00 | crowdstrike.com | original ↗ | #ai-security | #benchmark | #cybersecurity
3. Evaluating Pentesting Agents for the Real-World Part 2 research | 2026-07-22 00:00 | ethiack.com | original ↗ | #ai-security | #pentesting | #agentic-ai
4. Evaluating Pentesting Agents for the Real-World research | 2026-07-14 00:00 | ethiack.com | original ↗ | #ai-security | #pentesting | #agentic-ai
5. Inside the Top 1%: Engineering Tenzai’s AI Hacker to Compete with Elite Humans research | 2026-06-25 00:00 | tenzai.com | original ↗ | #ai-security | #benchmark | #ctf
6. Keep your Agents Under Control with agent-belt blog | 2026-05-19 13:17 | jfrog.com | original ↗ | #ai-security | #open-source | #mcp
7. Mythos for Offensive Security: XBOW's Evaluation research | 2026-05-12 12:00 | xbow.com | original ↗ | #ai-security | #benchmark | #reverse-engineering
8. DEF CON 33: Field Notes on AI Security, AI Red Teaming, and the Road Ahead blog | 2025-08-21 14:03 | hackerone.com | original ↗ | #ai-security | #def-con | #llm