Claude Opus 5: Frontier Intelligence at a Lower Cost per Task
Released 24 July 2026, Claude Opus 5 tops the Artificial Analysis Intelligence Index at 26% lower cost per task than Fable 5 — and beats it outright at less than half the price on high effort.
Anthropic released Claude Opus 5 on 24 July 2026. On the independent Artificial Analysis Intelligence Index it is now, narrowly, the most intelligent model evaluated — and it reaches that position at 26% lower cost per task than Claude Fable 5, the model it edges past.
That combination is the story. Opus 5 is not a marginal iteration. It is the point at which frontier-tier capability stopped carrying a frontier-tier price.
The numbers
At max effort, Claude Opus 5 scores 61 on the Artificial Analysis Intelligence Index. Claude Fable 5 sits at 60, GPT-5.6 Sol at 59, Kimi K3 at 57, and Claude Opus 4.8 at 56. The margin at the top is thin, but the cost difference is not: $2.03 per Intelligence Index task against $2.75 for Fable 5.
The clearer result is in agentic knowledge work, where the gap is not marginal at all:
- On GDPval-AA v2, Opus 5 at max effort scores 1861 Elo — more than 100 points ahead of both Claude Fable 5 and GPT-5.6 Sol
- On AA-Briefcase, Anthropic's strongest showing, it scores 1720 Elo, 146 points ahead of Fable 5
- Its xhigh (1693) and high (1606) settings also beat Fable 5, while costing considerably less
The cost curve underneath those numbers is what should interest anyone running budgets. Opus 5 at max effort costs $17.79 per AA-Briefcase task, about 20% below Fable 5 at $22.30. But at high effort it still outperforms Fable 5 — by 32 Elo — at $10.41 per task, less than half the price. The cheapest configuration that beats the previous frontier model is less than half its cost.
On coding, Opus 5 at xhigh running in Claude Code takes joint first place on the Artificial Analysis Coding Agent Index, including the highest recorded score on SWE-Atlas-QnA. On Terminal-Bench v2.1 it reaches 89% at max effort, roughly level with the leader, GPT-5.6 Sol. On Humanity's Last Exam it scores 53%, in line with Fable 5.
Pricing is unchanged from Opus 4.8: $5 per million input tokens and $25 per million output, with a 25% premium on cache writes and a 90% discount on cache reads. The context window is 1 million tokens.
What to watch out for
Three findings deserve more attention than the headline scores.
Hallucination rate went up. On AA-Omniscience, Opus 5 improves factual accuracy by 7 points over Opus 4.8 — but it also answers more often when uncertain, and its hallucination rate rises 14 points to 50%. A model that is both more accurate and more willing to guess is a genuinely mixed result. For any workload where a confident wrong answer is worse than an admitted gap, this needs an explicit verification layer. It is not a reason to avoid the model; it is a reason to design around it.
Presentation quality lags analysis. Opus 5's gains are driven by rubric pass rate and analytical quality — its Analytical Quality Elo of 2016 is nearly 300 points ahead of Fable 5. Its Presentation Elo of 1628 remains about 40 points behind GPT-5.6 Sol. If your output is a board deck, the analysis will be stronger and the formatting weaker.
Tasks take substantially longer. Opus 5's top three effort settings average more than 25 minutes per AA-Briefcase task — 36.2, 34.3, and 25.7 minutes for max, xhigh, and high. At max effort that is roughly 50% longer than Opus 4.8, driven mainly by turn count: 103 turns per task against 55 for Opus 4.8. The model is doing more work, which is why it scores better, but any system with a timeout tuned to Opus 4.8's behaviour will need revisiting.
Practical guidance
Do not default to max. The evidence points the other way: on this model, high and xhigh sit at a better point on the cost-quality curve for most work, and both already beat the previous frontier model. Max effort is for the cases where correctness genuinely outweighs cost, and it should be a deliberate choice.
Note also that thinking is now on by default — a request that omits the setting entirely will reason, where on Opus 4.8 it would not. Any workload that quietly relied on thinking being off will see both higher token consumption and, if its output limit was sized tightly around the answer alone, truncated responses. That is a small change with a large blast radius, and it is worth auditing before migration rather than after.
What this means for GCC enterprises
For eighteen months the practical constraint on deploying frontier models at scale has been cost per task, not capability. Opus 5 moves that line: at high effort it delivers better-than-Fable-5 results at under half Fable 5's price. Workloads that were shelved as economically unviable at the previous frontier are worth re-costing.
Three actions follow.
Re-run the business case on anything you rejected on cost grounds in the last year — the arithmetic has changed enough that some of those decisions no longer hold.
Sweep effort levels against your own evaluation set rather than inheriting a default. The spread between low and max on this model spans hundreds of Elo points and roughly eight times the token usage; the right setting is workload-specific and cannot be guessed.
Budget for verification, particularly given the higher hallucination rate. The cost saving from dropping to high effort more than covers a checking step, and on current evidence you want that step regardless of which frontier model you deploy.
The frontier is now genuinely contested — Anthropic and OpenAI traded the top position twice in a single month. For enterprises, the winner of that race matters far less than building systems that can switch between models as the ranking changes. Portability is the durable advantage. The leaderboard is not.
Ready to Apply This to Your Business?
Book a 30-minute strategy call. We'll take the thinking in this article and apply it directly to your workflows and business context.