Claude Opus 5 outperforms Fable 5 on agentic benchmark while cutting cost per task by 20%
New model release claims lead on Artificial Analysis’ AA-Briefcase benchmark and improves efficiency, but public evals show mixed results on standardized metrics.
4 sources · cross-referenced
- Anthropic released Claude Opus 5, positioning it as a leader on the Artificial Analysis Intelligence Index’s agentic knowledge work benchmark.
- Artificial Analysis reports Opus 5 outperforms Fable 5 by nearly 150 Elo on AA-Briefcase and reduces cost per task by 20%.
- Epoch AI’s ECI scores show Opus 5 at 159 versus Fable 5 at 161, with both models matching at 161 on software engineering benchmarks.
- Early user reports highlight strong coding and browser automation performance, though some note benchmark irregularities and small gains over Opus 4.8.
Anthropic released Claude Opus 5 on July 24, 2026, positioning it as the new leader on Artificial Analysis’ agentic knowledge work benchmark, AA-Briefcase. Artificial Analysis reported that Opus 5 outperforms Fable 5 by nearly 150 Elo while reducing cost per task by 20%.
Public benchmark results from Epoch AI indicate mixed outcomes. Epoch AI’s ECI scores place Opus 5 at 159, slightly below Fable 5’s 161, while both models match at 161 on software engineering benchmarks (SWE-ECI).
Early user reports and anecdotes suggest Opus 5 delivers stronger real-world performance in coding and tool-use tasks than aggregate benchmarks currently reflect. One practitioner reported a clear head-to-head win over Fable 5, especially with best-of-n sampling, and highlighted effective browser automation capabilities. Another noted that Opus 5 successfully automated browser tasks such as canceling a ChatGPT Pro subscription.
Some evaluators observed irregular benchmark behavior for Opus 5. One noted that on FrontierCode, medium effort outperformed high effort, contrary to the typical pattern where more effort yields better results. This raises questions about evaluation stability or task-specific tradeoffs in inference-time compute.
Beyond capability claims, distribution details emerged. Nous Research’s portal added access to Opus 5 with a 20% discount applied to all models, including Opus 5, reflecting broader availability and pricing strategy.
- Jul 25, 2026 · Simon Willison — everything
Anthropic unveils Claude Opus 5, positioning it as a proactive model with near-frontier performance at reduced cost
Trust79 - Jul 23, 2026 · OpenAI — News
OpenAI adds Health features to ChatGPT for eligible U.S. users
Trust76 - Jul 23, 2026 · Latent Space — swyx
Poolside AI releases Laguna S 2.1, a 118B-parameter MoE model outperforming larger open-weight models
Trust71