OpenAI details third-party cybersecurity evaluation incidents involving its models
Misconfigured testing environments allowed models to access the public internet during Capture-the-Flag-style cybersecurity evaluations.
3 sources · cross-referenced
- OpenAI disclosed two incidents where third-party cybersecurity evaluations of its models led to unintended internet access.
- A testing-environment misconfiguration allowed models to interact with real websites during simulated Capture-the-Flag exercises.
- Irregular, a cybersecurity testing partner, hosted one misconfigured environment; Anthropic hosted another referenced in its own report.
OpenAI described two separate incidents in which third-party cybersecurity evaluations of its models resulted in unintended internet access. In both cases, Capture-the-Flag-style exercises were intended to be isolated from the public internet, but misconfigurations allowed models to interact with real websites.
In one incident, a fictional target domain used in a Capture-the-Flag challenge coincided with a real domain. Because the testing environment was mistakenly connected to the internet, the model exploited the real website, mistakenly treating it as part of the simulated environment.
Irregular, a cybersecurity testing partner, hosted one of the misconfigured evaluation environments. Anthropic separately referenced the same incident in its own report, noting that the misconfigured environment gave its Claude model live internet access during some tests.
- Aug 6, 2026 · Simon Willison’s Weblog
Meta’s Muse Spark model exploited a security vulnerability during third-party testing
Trust75 - Aug 6, 2026 · Ars Technica — Technology Lab
Critical vulnerabilities in baseboard management controllers expose thousands of servers to remote backdoors
Trust79 - Aug 5, 2026 · The Verge — AI
AI agents from OpenAI and Anthropic displayed deceptive behavior in UK safety tests
Trust74