AI guardrails complicate offensive cybersecurity research, researchers say
Vetted access programs from OpenAI and Anthropic are creating inconsistent barriers for security researchers who rely on frontier models to probe systems for vulnerabilities.
1 source · cross-referenced
- OpenAI and Anthropic operate vetted programs that restrict how their models can be used for offensive cybersecurity research.
- Researchers report guardrails block legitimate vulnerability probing and create inconsistent access barriers.
- Some researchers bypass U.S. models for Chinese open-source alternatives to avoid restrictions.
- Export controls on Anthropic’s Mythos and Fable models were temporarily imposed and later lifted.
OpenAI and Anthropic have created vetted access programs—OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program—to allow approved researchers to use their models with fewer cybersecurity restrictions. These programs were designed to balance access with safety, but researchers say the guardrails are impeding their work.
Security researchers who probe systems for unknown vulnerabilities and develop exploits rely on AI tools to assist in tasks such as confirming bugs and reverse engineering. When guardrails block these activities, defenders lose a critical capability, according to Chris Anley, chief scientist at NCC Group. He described AI models as dual-use tools—essential for both offensive and defensive security work—comparing them to a hammer that is both a construction tool and a potential weapon.
Some researchers, including those at CrowdStrike subsidiary CrowdFense, avoid using frontier models for sensitive tasks like finding vulnerabilities or building exploits due to concerns about data leakage or exposure in future training runs. Instead, they use open-source models run locally, which come without guardrails but require more operational overhead.
A researcher at a smartphone-component manufacturer, speaking anonymously, said their employer’s tools are barely usable for security-related tasks because guardrails are too strict, blocking workflows when they detect potential security activity.
Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, reported that guardrails in frontier models are inconsistent and change frequently, forcing researchers to spend time negotiating with models rather than analyzing vulnerabilities. This inconsistency pushes some researchers toward Chinese open-source models like GLM, which can be run locally without restrictions.
Anthropic’s models Mythos and Fable were temporarily subject to U.S. export controls in June, which were later lifted—Fable 5 returned to general access on July 1, and Mythos 5 was reintroduced only to vetted U.S. organizations as part of a government review process.
- Jul 24, 2026 · TechCrunch — AI
AMD unveils Helios AI rack-scale system to compete with Nvidia’s Vera Rubin
Trust78 - Jul 23, 2026 · TechCrunch — AI
Anthropic updates Claude voice mode with more capable models and tool integrations
Trust79 - Jul 23, 2026 · TechCrunch — AI
AegisAI raises $36M Series A to counter AI-driven spear-phishing with agentic email defense
Trust79