Skip to content

Opus 5.5’s price cut shrinks against Fable exactly where AI agents spend the most

By 4 min read

According to Anthropic, their new model is 60% cheaper than Fable 5.1 on fresh tokens. On cached context, the highest cost in long agent runs, the gap narrows to 20%.

Anthropic released Claude Opus 5.5 today, Tuesday, September 22, pitching it as Fable 5.1-level performance on most work at 40% below the running cost of Opus 5. Most coverage has actually measured the new model against Opus 5. For companies running long agents, the more useful comparison is with Figure 5.1, and there the size of the discount depends on how the workload is built.

Four prices, four different discounts

On the Claude API pricing, Opus 5.5 costs $4 per million input tokens and $20 per million output. Cached context is read back at $0.20 per million tokens, and caching it costs $5 for a 5-minute cache or $8 for an hour. Fable 5.1 keeps its $10 input and $50 output rates, with 5-minute cache writes at $12.50 and cache reads at $0.25.

Line those up and Opus 5.5 undercuts Fable 5.1 by 60% on fresh input, output and cache writes, but by only 20% on cache reads. Anthropic says reused context accounts for most of what agentic and coding jobs cost.

The difference comes from how each model discounts cached tokens. Opus 5.5 bills cache reads at one-twentieth of its input rate. Fable 5.1 charges one-fortieth, while earlier Claude models used a one-tenth rate. Fable’s deeper cache discount erases most of Opus 5.5’s advantage on the tokens agents reuse most often.

How the math plays out

Take a hypothetical agent session that reads 20 million cached tokens, sends 2 million fresh input tokens and produces 500,000 output tokens. At list prices it runs about $22 on Opus 5.5, $32.50 on Opus 5 and $50 on Fable 5.1. That puts Opus 5.5 32% below Opus 5 and 56% below Fable.

Push the session further toward caching, with 100 million cached tokens, 1 million fresh input and 200,000 output, and the picture of course shifts considerably. Opus 5.5 costs about $28, Opus 5 costs $60 and Fable 5.1 costs $45. The saving over Opus 5 grows to 53%, while the saving over Fable drops to 38%.

The more an agent depends on cached context, the closer Opus 5.5’s advantage over Opus 5 moves toward 60%, and the closer its advantage over Fable falls toward 20%.

Anthropic also claims Opus 5.5 needs fewer tokens to finish a task, which would widen the gap.

Fast mode now undercuts Fable’s standard rate

A second pricing shift has gone largely unremarked. Opus 5.5’s fast mode, which generates output up to 2.5 times quicker, costs $8 per million input tokens and $40 per million output. Opus 5’s fast mode was priced at $10 and $50, the same as Fable 5.1’s regular rate.

Developers can now run Anthropic’s second-tier model at its accelerated setting for 20% less than its frontier model at normal speed. The option is limited to Anthropic’s own API. Fast mode isn’t offered on Amazon Bedrock, Claude Platform on AWS, Google Cloud or Microsoft Foundry. Customers who buy Claude through a cloud provider can’t reach the cheapest fast option.

Anthropic’s own tools are ready; everyone else has to catch up

Opus 5.5 locks its reasoning records (“thinking blocks”) to the conversation that produced them. Changing the system prompt, the tool list or earlier messages partway through breaks them. According to Anthropic’s migration guide, its own products (Claude Code, claude.ai, Claude Managed Agents and the Claude Agent SDK) already handle conversations in a way that avoids the problem.

On launch day, a GitHub issue on LibreChat, the open-source chat interface, requested Opus 5.5 support and warned that simply adding the model name wouldn’t be enough. It points to the always-on thinking, the rejection of forced tool calls and the conversation-bound reasoning blocks. It also asks maintainers to remove any thinking-off option and to re-test streaming, multi-step tool use, and switching between models.

The upshot: Opus 5.5 works cleanly from day one inside Anthropic’s products. I have tested it, and it serves answers fast. Anyone using it through a third-party front end may hit errors until the tool is updated.

Join the discussion

Be the first to join this discussion

Be respectful toward authors and fellow readers.

View authors
Portrait of Mercy Boluwatife

Written by

Mercy Boluwatife

View all stories