Anthropic’s Opus 5 shows measurable gains in resisting prompt injection attacks
On the IPI benchmark, Opus 5 reduces attacker success rates to 2.0% within 15 attempts and 0.2% on the first try, outperforming prior versions and other leading models.
1 source · cross-referenced
- Anthropic reports Opus 5 improves prompt-injection resistance on the IPI benchmark, cutting attacker success from 5.5% to 2.0% over 15 attempts and from 0.5% to 0.2% on the first attempt.
- Opus 5 outperforms Opus 4.8, Sonnet 5, Mythos 5, and all non-Claude models tested, including Muse Spark (16.5% within 15 attempts).
- GPT 5.6 variants show higher susceptibility: Sol at 20.0%, Terra at 30.4%, and Luna at 43.9% within 15 attempts.
- A single attempt against GPT 5.6 Sol succeeds 3.1% of the time, higher than Opus 5’s 2.0% after 15 attempts.
Anthropic’s latest system card for Opus 5 reports measurable gains in resisting prompt-injection attacks, citing results from the IPI benchmark. The model reduces the probability of an attacker succeeding within 15 attempts from 5.5% with Opus 4.8 to 2.0% with Opus 5, and from 0.5% to 0.2% on the first attempt.
Opus 5 also outperforms other Claude family models evaluated on the same benchmark. Against Sonnet 5, attacker success drops from 5.9% to 2.0% at k=15, and against Mythos 5, it falls from 2.6% to 0.2% on the first try.
Among non-Claude models, Muse Spark is the strongest performer with a 16.5% success rate within 15 attempts, more than eight times higher than Opus 5’s 2.0%. The most capable GPT 5.6 variant, Sol, shows a 20.0% success rate within 15 attempts, comparable to its predecessor GPT 5.5 (20.8%), and is ten times more likely to be successfully attacked than Opus 5.
GPT 5.6’s other variants fare worse: Terra at 30.4% and Luna at 43.9% within 15 attempts. Even on a single attempt, Sol achieves a 3.1% success rate, exceeding Opus 5’s 2.0% rate after 15 attempts.
The post notes that preventing prompt injection in the general case is considered impossible with current architectures, but these results suggest incremental hardening against specific attack patterns is achievable.
- Jul 31, 2026 · TechCrunch — AI
Anthropic reports three unauthorized system breaches by Claude models during internal security tests
Trust79 - Jul 31, 2026 · Simon Willison’s Weblog
Anthropic reports three real-world cybersecurity incidents during model evaluations
Trust79 - Jul 30, 2026 · Schneier on Security
Essay proposes ‘work vs. gym’ test to decide when AI assistance is appropriate
Trust76