An AI Model Just Topped a 'Best Hacker' Benchmark for $4.65 — Days After AI Agents Breached 395 Companies

Close-up of hands typing on a laptop keyboard with code on screen, representing AI-powered hacking tools

On September 10, 2026, security firm GreyNoise disclosed that an attacker had used swarms of AI agents to autonomously breach 395 organizations across 48 countries by exploiting flaws in PaperCut printing software. Days later, AI security research firm Enclave published benchmark results showing DeepSeek V4.1-Flash landing working code execution on all 11 vulnerable targets in its offensive-security benchmark — for a total cost of $4.65. Read together, the two stories are why Hacker News spent the week arguing about whether open-weight models are now arming attackers faster than defenders can keep up.

An 11-for-11 Score, For Under Five Dollars

Enclave's benchmark runs models against a mix of vulnerable and patched targets — apps like Grafana, Jenkins, and Nextcloud — and checks whether the model can turn a known flaw into actual code execution on the box. DeepSeek V4.1-Flash cleared all 11 vulnerable targets while leaving the four hardened targets untouched, avoiding false positives. The price tag is what made the result travel: $4.65 for the accepted runs, $5.14 once failed attempts are counted in.

The Mechanics Behind the Number

The cheapness wasn't a fluke of one lucky run. Enclave's own numbers show the model burned through 268.3 million input tokens — the vast majority of it cached — plus roughly 2 million output tokens across the full run, executing 2,349 separate Bash commands along the way. A typical successful compromise took a median of 4 minutes and 38 seconds, and the entire benchmark consumed about two hours and 38 minutes of active model time. Enclave also found that 5 of the 11 successful runs discovered exploit paths its own test environment hadn't anticipated — good enough findings that the firm went back and patched its benchmark to close the newly discovered routes.

The Breach That Made the Benchmark Feel Urgent

GreyNoise's report is what gave Enclave's numbers their edge. The firm said an attacker spent weeks running a swarm of AI agents against PaperCut NG/MF servers, exploiting two vulnerabilities tracked as CVE-2026-81578 and CVE-2026-82078. By the time GreyNoise finished counting, the campaign had touched at least 440 servers across those 395 organizations, with 280 organizations losing credentials, 147 having OS or domain secrets exposed, and 12 fully compromised with admin privileges.

Compromise at Machine Speed

What distinguished this campaign wasn't the vulnerability — it was the pace. GreyNoise says the attacker went from an empty workspace to remote code execution against a real victim in under four hours, and to first domain admin access two hours after that. Once the operation was fully running, it compromised 11 organizations in just 26 seconds. One high school reportedly went from initial access to full domain administrator control in seven minutes. GreyNoise attributes the speed to hundreds of AI agents working in parallel, reportedly built on tooling that included OpenAI's Codex and DeepSeek's models alongside conventional offensive-security tools, with a platform called Netlas used to generate target lists automatically. Education made up roughly half the victims, with the US, UK, France, Spain, and Canada hit hardest.

The Hacker News Fight

Enclave's post never claimed any link to the PaperCut campaign, and the two events involve different tooling and different targets. But landing in the same week, they collided on Hacker News into one argument about what a "best hacking model" claim actually proves.

Is the Benchmark Even Meaningful?

The loudest pushback was methodological: Enclave's post never compared DeepSeek V4.1-Flash against any competing model, leaving "best" as an unsupported claim. One commenter with hands-on reverse-engineering experience said GLM 5.3 turned up considerably more vulnerabilities for about $22 in a comprehensive run, while a $2 DeepSeek run found only a single issue — evidence, they argued, that the model is efficient at grabbing low-hanging fruit rather than doing deep analysis. Others noted that results swing heavily on deployment details like API provider, quantization level, and which agent harness is used, which can make identical models look wildly different across reports.

Cheap Enough to Parallelize

Where commenters agreed was on the economics. Developers reported running agent tasks on DeepSeek for under $1 apiece versus $3 to $15 for comparable jobs on other models — a gap wide enough that it becomes rational to run many cheap agents in parallel rather than one expensive one, even if each individual agent is less capable. Some cautioned that heavy reasoning-token usage can erode that advantage in practice. Others raised a sharper concern: DeepSeek's lighter content restrictions let it assist with security research that "American frontier models refused to help with," reviving the debate over whether open-weight models are shifting the offense-defense balance faster than policy can respond — with one thread comparing it to arms-control efforts like the Montreal Protocol, and another dismissing the safety framing as thinly veiled competition with Chinese labs.

What to Watch Next

Neither story alone would be this alarming — a benchmark is not a breach, and a breach built partly on AI tooling is not proof that any one model caused it. But together they mark a shift worth tracking: offensive AI capability is now cheap enough to buy for pocket change and, per GreyNoise, already being run at scale against real infrastructure. The open question isn't whether a model can pass a benchmark like Enclave's — it's whether defenders can patch, detect, and respond at the same machine speed attackers are now operating at.

-EditorZ

Photo by Towfiqu barbhuiya on Unsplash



 

OlderNewest

Post a Comment

Contact Form

Name

Email *

Message *