Skip to content
Tools · Jul 24, 2026

AI guardrails complicate offensive cybersecurity research, researchers say

Vetted access programs from OpenAI and Anthropic are creating inconsistent barriers for security researchers who rely on frontier models to probe systems for vulnerabilities.

Trust74
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • OpenAI and Anthropic operate vetted programs that restrict how their models can be used for offensive cybersecurity research.
  • Researchers report guardrails block legitimate vulnerability probing and create inconsistent access barriers.
  • Some researchers bypass U.S. models for Chinese open-source alternatives to avoid restrictions.
  • Export controls on Anthropic’s Mythos and Fable models were temporarily imposed and later lifted.

OpenAI and Anthropic have created vetted access programs—OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program—to allow approved researchers to use their models with fewer cybersecurity restrictions. These programs were designed to balance access with safety, but researchers say the guardrails are impeding their work.

Security researchers who probe systems for unknown vulnerabilities and develop exploits rely on AI tools to assist in tasks such as confirming bugs and reverse engineering. When guardrails block these activities, defenders lose a critical capability, according to Chris Anley, chief scientist at NCC Group. He described AI models as dual-use tools—essential for both offensive and defensive security work—comparing them to a hammer that is both a construction tool and a potential weapon.

Some researchers, including those at CrowdStrike subsidiary CrowdFense, avoid using frontier models for sensitive tasks like finding vulnerabilities or building exploits due to concerns about data leakage or exposure in future training runs. Instead, they use open-source models run locally, which come without guardrails but require more operational overhead.

A researcher at a smartphone-component manufacturer, speaking anonymously, said their employer’s tools are barely usable for security-related tasks because guardrails are too strict, blocking workflows when they detect potential security activity.

Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, reported that guardrails in frontier models are inconsistent and change frequently, forcing researchers to spend time negotiating with models rather than analyzing vulnerabilities. This inconsistency pushes some researchers toward Chinese open-source models like GLM, which can be run locally without restrictions.

Anthropic’s models Mythos and Fable were temporarily subject to U.S. export controls in June, which were later lifted—Fable 5 returned to general access on July 1, and Mythos 5 was reintroduced only to vetted U.S. organizations as part of a government review process.

Sources
  1. 01TechCrunch — AIHow AI guardrails are impeding the work of offensive cybersecurity researchers
Also on Tools

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.