Skip to content
Safety · Aug 6, 2026

OpenAI details third-party cybersecurity evaluation incidents involving its models

Misconfigured testing environments allowed models to access the public internet during Capture-the-Flag-style cybersecurity evaluations.

Trust79
HypeLow hype

3 sources · cross-referenced

ShareXLinkedInEmail
TL;DR
  • OpenAI disclosed two incidents where third-party cybersecurity evaluations of its models led to unintended internet access.
  • A testing-environment misconfiguration allowed models to interact with real websites during simulated Capture-the-Flag exercises.
  • Irregular, a cybersecurity testing partner, hosted one misconfigured environment; Anthropic hosted another referenced in its own report.

OpenAI described two separate incidents in which third-party cybersecurity evaluations of its models resulted in unintended internet access. In both cases, Capture-the-Flag-style exercises were intended to be isolated from the public internet, but misconfigurations allowed models to interact with real websites.

In one incident, a fictional target domain used in a Capture-the-Flag challenge coincided with a real domain. Because the testing environment was mistakenly connected to the internet, the model exploited the real website, mistakenly treating it as part of the simulated environment.

Irregular, a cybersecurity testing partner, hosted one of the misconfigured evaluation environments. Anthropic separately referenced the same incident in its own report, noting that the misconfigured environment gave its Claude model live internet access during some tests.

Sources
  1. 01Simon Willison’s WeblogThird-party cyber evaluations involving OpenAI models
  2. 02OpenAIThird-party cyber evaluations involving OpenAI models
  3. 03AnthropicInvestigating incidents in cybersecurity evaluations
Also on Safety

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.