Skip to content
winfunc
Research

Cyber guardrails tax the defender, not the attacker

Hugging Face had to fall back on an open-weight model to investigate an AI-run intrusion after Claude refused. That isn't a one-off. The way labs build cyber safeguards, the cost lands on defenders who need proof, not attackers who don't.

Mufeed VHPublished14 min read
  • ai-security
  • guardrails
  • open-models
  • vulnerability-research
  • policy
Cyber guardrails tax the defender, not the attacker

In July, an autonomous AI agent broke into Hugging Face and stayed for days. Hugging Face's responders later recovered about 17,600 of its actions between July 9 and July 13:

  • a sandbox escape
  • command-and-control staged on ordinary public web services
  • payloads chunked, XOR'd and compressed into dead drops
  • a pivot into one of their core clusters

"No human directed the individual steps," they wrote in their technical timeline.

The agent, driven by OpenAI models, was running an internal OpenAI cyber-capability evaluation built on the ExploitGym benchmark (OpenAI's account). Hugging Face believes the whole intrusion was, from the agent's point of view, an attempt to cheat: reach their production systems and steal the benchmark's reference solutions instead of solving the challenges. The evaluation had "deliberately disabled OpenAI's production safety classifiers" to measure the model's raw capability.

Reconstructing 17,600 actions by hand was impractical, so Hugging Face built an AI-assisted pipeline to rebuild the timeline and decode the payloads. The models they reached for first, Claude Opus and Fable, "refused a large part of that work: their safety guardrails treated reverse-engineering an exploit the same as launching one." The caption under their screenshot of the error says the rest: "Guardrails on Opus tripped every time we tried to analyze the attack logs."

So they stood up a quantized GLM-5.2, an open-weight model from Z.ai, on their own hardware and rerouted the entire pipeline through it. GLM-5.2 recovered the agent's chunk+XOR+compress scheme and the per-campaign key it had leaked across its own logs.

A commenter under the post wrote a summary I can't improve on:

closed model with disabled safeguards attacks -> other closed model with safeguards refuses to help defend -> open model saves the day

couldn't have written a better script

In July we published what happened when we ran our harness with GLM-5.1 and GLM-5.2 against NGINX. It found six bugs:

  • heap overflows in the rewrite engine, the stream scripting engine and two HTTP/2 upstream request builders
  • an HTTP/2 frame injection
  • an mTLS bypass that let a revoked client certificate through

They became five CVEs, and every F5 advisory credits Winfunc Research (write-up).

Here's my claim. The way cyber safeguards are built today, the cost lands on defenders, who need proof, and barely touches attackers, who don't.

Two disclosures first. I run Winfunc, an AI vulnerability research company, and looser safeguards are good for my business. Weigh what follows accordingly.

I'm also not a safety skeptic by temperament. Nine days after ChatGPT launched, I wrote about prompt injection, agents escaping their sandboxes, and jailbreaks "gaining code execution with the exploit being just plain english" (Security in the age of LLMs). I think the offensive capability is real. That's why I care where the guardrail sits.

Where the guardrail sits

I'll mostly quote Anthropic, because its documents spell out the policy in unusual detail. The pattern isn't unique to them. OpenAI gates its cyber-permissive models behind Trusted Access for Cyber and Daybreak. Google is rolling out Gemini 4 Argon to vetted defenders first.

Fable 5's classifiers sort cyber requests into four buckets: prohibited, high-risk dual use, low-risk dual use, and benign (Anthropic). The high-risk bucket is blocked "until we have better controls to limit access to known good actors." It reads like a security team's job description:

  • "Hacking, penetration testing, red teaming, and bug bounties"
  • "Exploit development and weaponization (including zero-click and memory-corruption work)"
  • "Virtual machine or container escapes"
  • "High-uplift vulnerability finding: vulnerabilities that are not easily found by other widely available models"

On exploits, the policy is blunt: "we block the automatic generation of exploits." Vulnerability finding is allowed in principle, but only where "many widely available models in the industry are capable of finding that vulnerability."

Then there's the safety margin. Classifiers can be jailbroken, so Anthropic deliberately tunes them to fire on requests it knows are probably benign. A request "has to look very clearly safe to avoid triggering the classifier." For Fable 5 the margin is "much larger than in any prior launch," "meaning that many more benign requests would be blocked" (Anthropic).

Blocked requests get answered by an older model, Claude Opus 4.8. Opus 5.5, released September 22, works the same way: "most cybersecurity tasks will be re-routed to Opus 4.8" (Anthropic).

Put that together and the public frontier model will help you find a bug, as long as other models could have found it too, and is designed not to help you prove it.

Proof is the expensive part

Finding bugs stopped being the bottleneck this year. Anthropic said it in May: "Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it's limited by how quickly we can verify, disclose, and patch the large numbers of vulnerabilities found by AI" (Glasswing update). On October 1, Google stopped accepting product vulnerability reports to its open-source bug bounty after a flood of invalid AI-generated submissions (TechCrunch).

What separates a real finding from a confident hallucination is a reproducer. In March I asked maintainers to demand one, PoC||GTFO (write-up), and our pipeline is built around that filter. Most model hypotheses are wrong. The step that kills them is the one that turns an idea into a harness and runs it against a real build. Without a working reproducer, the finding dies there.

Chris Anley, chief scientist at NCC Group, made the same point to TechCrunch. Asking a model to try to exploit a bug is a key step in confirming it's worth fixing, so a guardrail that refuses that step hurts defenders (TechCrunch). His description of the dual-use problem is hard to beat: "'fix this code' as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base... the two can't really be unpicked."

Look at what made our NGINX reports land:

  • For the rewrite bug, an AddressSanitizer trace of a copy writing past a 4,096-byte request-pool allocation.
  • For the stream scripting bug, NGINX copying 12,000 bytes into a one-byte allocation.

Nobody at F5 had to take a model's word for anything.

Fable 5's policy files "memory-corruption work" under exploit development, in the blocked column. A reproducer that drives NGINX into an out-of-bounds write is hard to describe as anything else.

The safeguard is on the wrong side of the proof.

For the attacker, the PoC is the attack

Now look at the same step from the other side.

A defender needs proof because a defender has to convince someone: a maintainer, a vendor PSIRT, a change advisory board. An attacker has to convince no one. The target is the oracle. Fire the exploit attempt. If it lands, that's your proof. If it doesn't, fire the next one.

The Hugging Face intrusion shows how little being wrong costs offense. Most of the agent's actions "went nowhere," Hugging Face wrote. "Together, however, they produced enough coverage to find a viable chain across several independent systems." Black Ice Labs put the asymmetry in a title: Attackers don't have to be right. Defenders do.

A refusal on exploit generation removes the step defenders can't skip and attackers never needed.

Their model never asks

And when attackers do want help, they have it.

In September, NIST's CAISI called Z.ai's GLM-5.3 "the most cyber-capable open-weight model released to date" and put it about four months behind the US frontier (CAISI). In Anthropic's own evaluation, GLM-5.3 built working end-to-end exploits for Chrome V8 bugs in 50 of 410 attempts. Claude Mythos Preview managed 56 (Anthropic).

GLM-5.3's safeguards held against a bare malicious request and folded against standard tricks:

  • 64% compliance with a false cover story
  • 92% with prefilled reasoning
  • 100% once abliterated (refusals removed from the weights)

The UK AI Security Institute measured the open-weight lag on cyber at four to seven months in July, down from six to ten through most of 2025 (AISI). The smaller GLM-5.3-Flash runs locally on roughly $5,000 to $15,000 of hardware, and abliterated builds are already on Hugging Face (jyn).

One detail in the CAISI comparison is easy to miss. As Anthropic notes, the US models in it were "tested with cyber safeguards disabled when applicable," and the US frontier "includes models released only to vetted users." The four-month gap is measured against models most defenders aren't allowed to use. For a defender on a public tier, the honest comparison is an abliterated GLM-5.3 against whatever the classifier lets through.

Exploit shops already work this way. Paolo Stagno, CTO of Crowdfense, which develops and sells zero-days to government agencies, told TechCrunch that AI companies "essentially treat customers like children who need babysitting." For vulnerability finding and exploit work, his team runs open models locally, because sending that work to a cloud model risks leaking vulnerability data. Hugging Face made the same move for incident response, "with the added benefit of keeping the attacker data on-prem."

A jailbreak with zero uplift

If you want this whole argument in one incident, it's Fable 5's launch.

Fable 5 and Mythos 5 shipped on June 9. They share one underlying model: Fable 5 got the strong safeguards, and Mythos 5, with fewer, went only to Glasswing partners. Three days later the US government applied export controls to both.

According to Anthropic, the controls followed a report in which Amazon researchers had prompted Fable 5 into identifying a number of vulnerabilities and, in one case, writing code that demonstrated an exploit. Anthropic had no way to verify nationality in real time, so it suspended both models for every user (Anthropic). TechCrunch has argued the controls were never really about the jailbreak. Either way, the jailbreak is what got fixed.

Then Anthropic investigated. Many less capable models, including Opus 4.8, GPT-5.5 and Kimi K2.7, found the same vulnerabilities. For the exploit demonstration, "every model we tested could produce the same demonstration as Fable 5," all the way down to Claude Haiku 4.5. Anthropic concluded that the bypass "did not expose any unique Mythos-level cyber capabilities" and "only involved routine defensive cybersecurity work."

The fix was a new classifier that blocks the reported technique more than 99% of the time. In Anthropic's words, it "also comes at the cost of flagging benign requests more often during routine coding and debugging tasks."

So a bypass that gave attackers nothing they couldn't get from Haiku was answered with a classifier that blocks more routine work for everyone else.

So who pays?

Not the Glasswing partners. Amazon Web Services, Apple, Google, JPMorganChase, Microsoft and Nvidia got Mythos Preview in April (Foreign Policy). Anthropic says partners verified at least 129,000 vulnerabilities between April and July (Anthropic). That work matters, and it happened because they got in early.

Everyone else applies. On October 6, Anthropic reorganized its Cyber Verification Program into three tiers:

  • Defense Access covers incident response, malware reverse engineering and "analyzing and validating vulnerabilities." Individual researchers with a track record can qualify.
  • Red Team Access, for authorized pentesting, takes a few weeks to review and "is for organizations only; individual researchers are not eligible."
  • A third, Specialized Access, is reserved for organizations testing critical systems like power grids and telecom networks.

Applying requires identity verification. As of late September, Anthropic's help center also said organizations on zero data retention weren't eligible (Claude Help Center). If you hold unpatched vulnerability data and won't let a vendor retain it, you pick between retention and refusals.

I got into this field as a teenager, sending reports to the bug bounty programs of Google, Mastercard, Okta, etc. Bug bounties are on Fable 5's high-risk list. By name. The version of me starting out today would need a track record to get the model that helps build a track record, and the red-team tier wouldn't take him at all. The abliterated GLM on Hugging Face doesn't ask for anything.

Students hit the same wall. Researchers at Scale AI took 2,390 real-world requests from the National Collegiate Cyber Defense Competition, a blue-team event (Defensive Refusal Bias). Models refused defensive requests containing security-sensitive keywords at 2.72 times the rate of equivalent neutral ones. The worst rates landed on the most operational work: system hardening at 43.8% and malware analysis at 34.3%. When users said they were authorized, refusals went up. The models, the authors write, "interpret justifications as adversarial rather than exculpatory."

It's not hard to see why. The best-documented AI-orchestrated espionage campaign, GTG-1002, got Claude to cooperate in 2025 partly by telling it that it "was an employee of a legitimate cybersecurity firm, and was being used in defensive testing" (Anthropic). Anthropic has since hardened against that; in its GLM-5.3 tests, Claude Opus 5 went along with none of the cover-story attacks. But to a classifier, an honest scope statement and a pretext are the same string. Defenders now pay for GTG-1002's pretext every time they tell the truth.

The refusal you never see

A refusal is at least honest. It tells you to go elsewhere.

James Kettle at PortSwigger described something worse this week (PortSwigger). Throwing swarms of agents at open-ended research, he found the model "would try to complete its objective, but use the wiggle-room in the prompt to steer in a direction that actually sabotaged its performance." It drifted toward low-impact behavior that's easy to observe.

He first noticed it on OpenAI's daybreak-blue, powered by gpt-5.6-sol, the tier OpenAI runs for verified defenders. In the eval he built afterwards, "no refusals were received from any models." His takeaway: "just because a model agrees to do what you ask, doesn't mean it's really cooperating." He suspects alignment and plans to test an abliterated model next, so treat the cause as open.

This hurts agents more than people. The Scale AI authors point out that autonomous defensive agents "cannot rephrase refused queries or retry." In an unattended pipeline, a refusal doesn't look like an error. It looks like a hypothesis that didn't pan out, and nobody goes back to check.

The best case for the guardrails

The other side deserves its strongest version:

  • The capability jump is real. On Anthropic's V8 benchmark, Opus 4.6 and GLM-5.2 sit at or near zero, while Mythos Preview and GLM-5.3 reach 14% and 12%.
  • Gating bought defenders a head start. Those 129,000 verified vulnerabilities came out of it. Mozilla fixed 271 bugs in Firefox 150 while testing Mythos Preview, more than ten times what it found in Firefox 148 with Opus 4.6 (Anthropic).
  • Claude's safeguards work. It held up against the pretext and prefill attacks that broke GLM-5.3.
  • Hugging Face hit false positives, not the policy. Anthropic's own policy lists incident response, log analysis and malware reverse engineering as benign.
  • Anthropic isn't against open weights or wider access. It has rejected calls to ban open-weight models. Its GLM-5.3 post says the model "underscores the urgency of expanding access to advanced frontier models to a broader set of entities to empower cyber defenders."

I agree with more of that than this post might suggest. I don't think anonymous accounts should get push-button V8 chains either.

But a head start depreciates. Anthropic's GLM-5.3 post concedes as much in one line: "But those models have now arrived." Once a comparable model is a download away, a refusal stops buying time. It only sets the price defenders pay.

Put the guardrail where the harm is

Here's what I'd change.

Gate the weapon, not the proof. A crash input, a sanitizer trace, a failing test, a minimal trigger: that's what a maintainer needs to ship a fix. Mitigation bypasses, reliable chains, payloads, delivery and C2 are what turn a heap overflow into a shell, and much of that list is already prohibited outright. Anthropic's own benchmarks say the new, dangerous capability is end-to-end exploitation, so draw the line there. None of our NGINX reports needed a weapon. The gRPC report says it plainly: "It does not claim reliable remote code execution." The ASLR bypass F5's advisory mentions was never demonstrated, and it didn't need to be.

Score refusals the way you score jailbreaks. Anthropic's draft jailbreak severity framework grades a bypass by "the capabilities the jailbreak unblocks for attackers that they would not otherwise have had" (Anthropic). Run the same counterfactual before blocking a defender. Fable 5's vulnerability-finding policy already goes halfway: if widely available models can find a bug, Fable may find it too. Extend that to reproducers and incident response. The June investigation is a worked example of a block with zero counterfactual value.

Verify people with the credentials security already trusts. CVE credits, vendor advisories and bounty histories are public and checkable; advisories name the researcher. Let individuals with that kind of record into the red-team tiers, and stop making data retention the price of trust.

Publish the defender false-positive rate. Anthropic says Fable 5's safeguards trigger "in less than 5% of sessions" on average (Anthropic). Averaged across every user, that says little about sessions that are actually security work, which is presumably where most triggers land. Publish refusal rates on a defensive benchmark like the NCCDC set. Then Cisco Talos's advice to track model refusal rates becomes something a SOC can plan around (Talos).

Keep the logs, drop the reflex. Anthropic caught GTG-1002 the way security teams catch most things: it "detected suspicious activity," spent ten days mapping the operation, banned accounts and notified affected organizations. The pretext had already gotten past the model. Telemetry is the one durable advantage a hosted model has over a download. Every unnecessary refusal pushes defenders somewhere with no logs at all, like Hugging Face's on-prem GLM box. The exploit brokers got there first.

The bet

When we started the company, the founding bet was that AI would make attacking software cheap long before it made defending software cheap. I'd like to lose that bet.

Right now, the public frontier model will help you find a bug other models could find, then decline to help you prove it. The attacker's model does both, and it never needed the proof. The safeguards, as built today, are helping me win.

Written by

Mufeed VH

Co-founder and CEO at Winfunc.

Continue reading