Four frontier AI models couldn't exploit a single textbook CVE. Here's what that actually means.
A live-target benchmark ran Claude Opus 4.8, Sonnet 4.6, Devstral 2 123B, and Amazon Nova Pro against five well-known CVEs. 100 runs, zero exploits landed.
We ran four frontier AI models against live Vulhub Docker containers with five well-documented CVEs. No CTF flags, no synthetic environments: real unpatched targets, real pass criteria tied to confirmed RCE or unauthorized data access.
100 runs. Zero exploits landed. Every model: 0%.
Before you file this under “AI security is a non-issue,” read the failure profiles. The headline number is the least interesting part.
The benchmark
The campaign is called security-cve-v1. It ran in late June 2026 against five CVEs chosen specifically because they shouldn’t be hard:
| CVE | App | Class | CVSS |
|---|---|---|---|
| CVE-2021-44228 | Log4j 2.14.1 | Log4Shell (JNDI injection) | 10.0 |
| CVE-2022-22963 | Spring Cloud 3.1.6 | SpEL RCE via routing expression | 9.8 |
| CVE-2017-5638 | Apache Struts 2.3.5 | OGNL injection via Content-Type | 10.0 |
| CVE-2022-46169 | Cacti 1.2.22 | Unauthenticated command injection | 9.8 |
| CVE-2021-41773 | Apache HTTPD 2.4.49 | Path traversal + RCE | 9.8 |
These are the CVEs that appear in every OSCP prep guide, every security certification course, every “learn pentesting” YouTube channel from 2022 onward. The PoC code for all five has been in frontier LLM training data for years. If autonomous AI exploitation was going to show up anywhere, it was going to show up here.
Each model got a shell tool, access to the source code, and a pass criterion. No human in the loop. The harness scored on actual exploitation outcome: not “explained the vulnerability correctly,” not “generated plausible-looking payload.” Confirmed RCE or data access, or it’s a fail.
Four models, three failure modes
Here’s how it broke down:
Claude Opus 4.8 ran 25 times, spent an average of $2.16 per run ($53.93 total), averaged 13-25 tool calls per run, and failed all 25. On Log4Shell specifically, it averaged 25.2 tool calls per run at $4.02 average cost. Four runs burned through the entire $5 budget still trying. The failure classification was wrong_answer 84% of the time: the model submitted an attempt, the harness rejected it.
Claude Sonnet 4.6 ran 25 times and failed all 25 with wrong_answer every time. It was the most active model in the dataset, averaging 22 to 33 tool calls per run across CVEs, higher than Opus on every single task. Log4Shell consumed 1.67 million input tokens across five runs. Cost tracking has a known bug in this campaign (recorded as $0 despite estimated ~$20 in token consumption). Sonnet tried harder than any other model in the cohort. It still landed nothing.
Devstral 2 123B ran 25 times, spent $0.09 total ($0.004 per run), averaged 2-6 tool calls per run, finished each in 13-20 seconds, and failed all 25. Log4Shell, the same task that pushed Opus to budget exhaustion, averaged 2.0 tool calls from Devstral. wrong_answer on all 25 runs.
Amazon Nova Pro ran 25 times and failed all 25 at turn zero. Latency under one second. Zero tokens output. The model’s content filter classified the task as out-of-scope. Nova Pro never attempted any CVE.
What this tells defenders
The 0% headline is reassuring in one narrow sense: as of mid-2026, you’re probably not facing autonomous AI agents that can independently close the loop on textbook CVEs against live targets in standard API usage.
The key phrase is “standard API usage with a fixed budget and no human in the loop.”
The failure modes for Opus and Sonnet are not “the model refused” or “the model didn’t know what to do.” Both models identified the vulnerable code paths, assembled payloads, and iterated. The failure is in execution over many steps: maintaining a coherent exploitation chain across 25-33 tool calls against live infrastructure. The models produce sophisticated-looking attempts that don’t satisfy the pass criterion on a live target.
That’s a different threat model than “AI can’t exploit anything.” What it suggests for defenders:
The gap isn’t in recognition. These models understand the CVEs. Detection and triage workflows that rely on AI to identify or explain vulnerabilities are not in question. That’s not what failed here.
The gap is in agentic execution: maintaining state, interpreting live responses, adjusting payloads across many steps without human correction. That’s the current capability ceiling for autonomous exploitation.
Nova Pro’s content filter is real, but treating “the model refused” as a general security property of LLMs would be a mistake. Opus and Sonnet ran the same prompts without refusal. Nova Pro’s behavior is a policy choice, not a capability ceiling.
What this tells security tool builders
If you’re building LLM-assisted security tooling, the cost data from this campaign is worth studying even though everything failed.
Opus at $2.16 per run means running a single model across five CVEs costs $11 per attempt at current pricing. Log4Shell alone costs $4.02 per run on average. Devstral 2 costs $0.004 per run, three orders of magnitude cheaper, with the same null result.
For red-team tooling where you’re feeding models vulnerability descriptions and asking them to generate exploitation attempts for human review (not autonomous execution), the cost curve looks very different from this benchmark. This campaign measures fully autonomous execution, which is where the expensive iterations happen. Human-directed workflows with AI-generated draft payloads would consume far fewer tokens per useful output.
The Devstral data point is interesting for anyone considering open-weight models in security tooling: it’s cheaper, but in this benchmark it also barely tried. At 2 tool calls per Log4Shell run, you’re not getting the same depth of exploration as a Claude model. Whether that matters for your specific workflow depends on what you’re asking the model to do.
What the open-weights comparison doesn’t settle
The expected finding going in was that Devstral 2 (code-focused, open-weight, less safety-constrained) would outperform Claude models by executing exploits the Claude models might refuse.
That’s not what happened. Devstral scored the same as Opus and Sonnet but barely tried. Claude models engaged with the prompts without refusal and still failed. The data doesn’t support the “alignment tax” narrative for this benchmark. It supports a capability-gap narrative, with Nova Pro’s content filter as a separate, clearly labeled data point.
This doesn’t settle the open-weights vs. safety-trained question for security research more broadly. It settles it for this specific setup: five textbook CVEs, live targets, fixed token budget, fully autonomous execution. Devstral’s low engagement might reflect architectural limitations in its agentic loop, a different failure strategy, or something about how it handles offensive security prompts at the task level. We can’t distinguish those from this data alone.
What we’d want to test next
A few gaps this campaign doesn’t close:
Human-in-the-loop setups. If a security researcher is directing the model step-by-step (reviewing each tool call, correcting mistakes, re-prompting on response bytes), the execution gap narrows significantly. This benchmark is specifically designed to isolate autonomous capability. It doesn’t measure assisted capability.
Different agentic scaffolding. The campaign used standard API calls with a shell tool. Purpose-built exploitation frameworks with specialized prompting, structured memory, and multi-agent pipelines might perform differently. We don’t have data on that.
Fine-tuned models. A model specifically trained on exploitation workflows, with examples of successful CVE exploitation in the training set, is a different animal from a general-purpose frontier model. Security-focused fine-tunes weren’t in this cohort.
The full analysis (scoreboard breakdowns, per-CVE cost and latency data, failure mode histograms, and cross-model behavioral comparisons) is at modelbattles.com/articles/security-cve-v1-2026.
Related reading
- Best CVE intelligence tools for bug bounty hunters — how researchers track and triage new CVEs
- June 2026 CVE opportunity analysis — high-bounty vulnerabilities worth hunting this month