Skip to content
Models · Jul 25, 2026

Claude Opus 5 outperforms Fable 5 on agentic benchmark while cutting cost per task by 20%

New model release claims lead on Artificial Analysis’ AA-Briefcase benchmark and improves efficiency, but public evals show mixed results on standardized metrics.

Trust71
HypeLow hype

4 sources · cross-referenced

ShareXLinkedInEmail
TL;DR
  • Anthropic released Claude Opus 5, positioning it as a leader on the Artificial Analysis Intelligence Index’s agentic knowledge work benchmark.
  • Artificial Analysis reports Opus 5 outperforms Fable 5 by nearly 150 Elo on AA-Briefcase and reduces cost per task by 20%.
  • Epoch AI’s ECI scores show Opus 5 at 159 versus Fable 5 at 161, with both models matching at 161 on software engineering benchmarks.
  • Early user reports highlight strong coding and browser automation performance, though some note benchmark irregularities and small gains over Opus 4.8.

Anthropic released Claude Opus 5 on July 24, 2026, positioning it as the new leader on Artificial Analysis’ agentic knowledge work benchmark, AA-Briefcase. Artificial Analysis reported that Opus 5 outperforms Fable 5 by nearly 150 Elo while reducing cost per task by 20%.

Public benchmark results from Epoch AI indicate mixed outcomes. Epoch AI’s ECI scores place Opus 5 at 159, slightly below Fable 5’s 161, while both models match at 161 on software engineering benchmarks (SWE-ECI).

Early user reports and anecdotes suggest Opus 5 delivers stronger real-world performance in coding and tool-use tasks than aggregate benchmarks currently reflect. One practitioner reported a clear head-to-head win over Fable 5, especially with best-of-n sampling, and highlighted effective browser automation capabilities. Another noted that Opus 5 successfully automated browser tasks such as canceling a ChatGPT Pro subscription.

Some evaluators observed irregular benchmark behavior for Opus 5. One noted that on FrontierCode, medium effort outperformed high effort, contrary to the typical pattern where more effort yields better results. This raises questions about evaluation stability or task-specific tradeoffs in inference-time compute.

Beyond capability claims, distribution details emerged. Nous Research’s portal added access to Opus 5 with a 20% discount applied to all models, including Opus 5, reflecting broader availability and pricing strategy.

Sources
  1. 01Latent Space — swyx[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable)
  2. 02Artificial AnalysisClaude Opus 5 is the new leader on our agentic knowledge work benchmark, AA-Briefcase, outperforming Claude Fable 5 by nearly 150 Elo while reducing Cost per Task by 20%
  3. 03Epoch AI ResearchEpoch reported that Claude Opus 5 achieves an ECI of 159, slightly below Fable 5’s value of 161, while matching Fable 5 on SWE-ECI at 161 on software engineering benchmarks
  4. 04Nous ResearchNous Research’s portal added access to the model, with a tweet saying users could directly use Opus 5 through Nous Portal and that a 20% discount applied to all models including Opus 5
Also on Models

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.