Skip to content
Research · Aug 19, 2026

Preprint quantifies cost and accuracy impact of explicit reasoning-effort terms in model API contracts

Registered experiment on Sonnet 5 finds explicit high-effort contracts raise mean cost per call by about one cent while yielding no statistically detectable accuracy gain on AIME 2026 problems.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • A preregistered study on Sonnet 5 evaluated the effect of an explicit high-reasoning-effort term in the API contract versus the same model with effort omitted.
  • Mean delivered cost per call increased by $0.01031 under the explicit-high contract, with a 95% confidence interval of [$0.00204, $0.01974].
  • Accuracy difference was +0.0133 points [-0.0267, +0.0467]; the interval allows for up to a 4.67 percentage-point gain the design could not rule out.
  • Cost per correct answer was $0.08665 under the high-effort contract and $0.07662 under the omitted contract, per registered point estimates.

Researchers conducted a preregistered, paired-contrast experiment comparing Sonnet 5 under two contract conditions: one with an explicit high-reasoning-effort term and one with the term omitted. The study used 30 AIME 2026 items and five API calls per item, for a total of 300 paid attempts.

Every paid attempt was assigned a frozen terminal category, and the inference protocol resampled items while retaining repeated calls to the same item. The primary outcome was mean delivered cost per call, which was $0.01031 higher under the explicit-high contract than under the omitted contract, with a 95% confidence interval of [$0.00204, $0.01974].

The corresponding accuracy contrast was +0.0133 points, with a 95% interval of [-0.0267, +0.0467]. The authors report that no statistically detectable accuracy difference was observed, and the interval permits a potential gain of up to 4.67 percentage points that the design could not rule out.

Registered point estimates for cost per correct answer were $0.08665 under the high-effort contract and $0.07662 under the omitted contract. The authors also document model-specific omission semantics via a dated contract census, Models-API metadata, and preregistered raw-response probes, noting that claims remained at documentation grade when raw structure was indeterminate.

The request registry, parser, terminal taxonomy, statistical plan, and analysis pipeline were frozen before outcomes were examined, bounding the claims to the specific model, task, and collection date studied.

Sources
  1. 01arXiv cs.AIThe Price of Thinking: Reasoning Effort as a Model-Specific API Contract
Also on Research

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.