Claude Opus 5

Claude Opus 5: What a Flagship Costs, and What It Buys

On OrcaRouter we host the Claude Opus 5 model — Anthropic’s top-of-line flagship, shipped on July 24, 2026 at a $5 / $25 rate card with a 1M-token context window and a 128K output ceiling, a card whose pricing makes sense only if the model is doing work that outlasts a single call. It scores 51 on Artificial Analysis’s Intelligence Index at the max reasoning setting, and its production traffic in the trailing seven days is the highest in the Anthropic family at 23.0 million tokens. It is served at list price on the same key as every other Anthropic model, with the live telemetry on the linked page.

The number that deserves a second look is the traffic. 23.0 million tokens a week is the most of any Anthropic card we serve, and it is running through a $5 / $25 endpoint — the second most expensive in the family. Someone is paying flagship prices in volume, and that tells you what they are building.

Who pays $5 / $25 in volume

Two workloads justify an expensive endpoint used heavily. The first is agent loops: a task that reasons for many turns, where the input is cheap relative to the number of reasoning steps the model does inside its own context. The pricing structure — $5 input against $25 output — is visible proof of that assumption: the card makes reading a lot cheap-ish and writing expensive, which only pays off when the model writes a little and decides a lot.

The second is high-stakes single tasks where correctness outranks cost: a migration analysis, a contract review, a design decision that a cheaper model would need to be re-run ten times to match. For those, the price per million tokens is noise next to the price of a wrong output. The 23.0 million weekly tokens are mostly one of these two shapes.

The 128K output ceiling is part of the same design: the card is built to finish long jobs in one pass, not to stream short ones. Teams running agent loops on this endpoint batch their decisions inside a single context, read a lot, decide a lot, and write the conclusion once — which is exactly the shape the $5 / $25 meter rewards. A workload that is mostly output does not need this card; a workload that is mostly deciding is what the meter is for.

The cost question, in numbers

modelcontextinput / outputAA index (config)per-task costp50 TTFT
Claude Opus 51M$5 / $2551 (max)$5.864.66 s
Claude Sonnet 5.51M$2 / $1056 (max with fallback)$5.464.14 s
Claude Fable 5.11M$10 / $5053 (max with fallback)$7.634.04 s

Rate cards are what Anthropic lists; the index and per-task costs are Artificial Analysis’s live readings; the TTFT column is production telemetry from the linked page. Read the table carefully and the flagship’s case gets uncomfortable: the mid-tier Sonnet 5.5 outscores it on the index at 56 (max with fallback), and the Fable card outscores it at 53 while costing double. The flagship’s defense is not the score — it is everything the index does not measure.

The ledger column worth staring at is per-task cost: $5.86 for the flagship against $5.46 for the mid-tier on the same standard task set. The list-price gap is two and a half times, yet the task-level bills nearly converge, because the flagship reads and writes more tokens to finish the same job. That is the honest arithmetic of agent pricing: you are not paying for answers, you are paying for the model’s efficiency at closing the loop — and the two columns measure different things.

What the index cannot capture

Agent work and long-horizon tasks are poorly represented in the standard task set. The evaluator measures single-turn reasoning quality; the flagship’s entire purpose is the opposite — sustained execution, tool discipline across a long session, the ability to hold a plan for twenty steps without drifting. None of that shows up in a 51. The per-task cost figure of $5.86 is likewise a single-turn number; in an agent loop the effective cost is the model’s efficiency at completing the loop, which no static task set models.

That is the honest way to read this model’s numbers: the independent score is real but measures the wrong game, and the production telemetry is the signal that the right game is being played. 23.0 million tokens a week on the family’s most expensive card is the market’s own verdict that the agent-priced flagship earns its keep where it matters.

Reading the real telemetry

The linked page shows the endpoint’s own seven-day window: a p50 first token of 4.66 seconds and 23.0 million tokens of actual traffic. Note what is absent from that page: no synthetic benchmark, no vendor claim, just the model’s own production numbers on the same key you would use. For a team weighing the flagship, that page answers the two questions that actually decide a purchase — what does it cost at list price, and how does it behave under real load — with one screenshot.

The 4.66-second p50 is also a quiet cost. Every agent loop multiplies latency by turn count, so a flagship that takes two seconds longer than a mid-tier card on first token pays that difference across every loop iteration. The card’s economics assume the loop completes in fewer turns; if it does not, the $5 / $25 price gets painful. That is the real evaluation: measure your own loop completion, not the score.

The honest caution sits in the same window: 23.0 million tokens is real traffic, but it is traffic on a card new enough to still be ramping. Early adopters of a flagship are exactly the teams whose workloads justify it, which makes the volume partly selection — the people running this card are the people who chose it on purpose. The number still answers the question that matters: at $5 / $25, enough teams chose it to make the meter the family’s busiest.

The takeaway

Claude Opus 5 is Anthropic’s most expensive general card, priced at $5 / $25 for agent-shaped workloads that outlast a single call. Its 51 (max) index score is honest but no longer the family’s best — Sonnet 5.5 leads at 56 — and its measured p50 first token of 4.66 seconds is not the fastest. What it does have is 23.0 million tokens of weekly production traffic, the most in the family, which is the market’s answer to the question of whether the flagship earns its card.

Choose it when your work is loops and long-horizon execution and you can measure loop completion against the price. Choose the mid-tier card when a single-turn task is the workload, because the score and the cost both point there. The telemetry page linked above lets you make that call against your own numbers rather than someone else’s benchmark.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *