Anthropic's Claude Fable 5.1 and Claude Mythos 5.1 launched on September 1, 2026, and most of the coverage focused on the obvious stuff — benchmark bumps, a new model name to memorize. The detail that actually matters if you run Claude Code sessions all day got buried: cache read pricing dropped 75%, from $1.00 to $0.25 per million tokens. That's not a marketing number. It's the kind of change that quietly reshapes what a long agentic session costs you by the end of the month.
What actually shipped
Fable 5.1 and Mythos 5.1 are point releases on top of June's Fable 5 and Mythos 5 — same $10-per-million-input / $50-per-million-output standard pricing, same general shape, but with meaningful gains on agentic benchmarks and, notably, far fewer safety-driven interruptions during autonomous work. Anthropic says Claude Code sessions on 5.1 see around 60% fewer interventions than on the previous version, which in practice means fewer moments where a long-running task stalls out waiting for a safety check that didn't need to trigger.
But the pricing change is the part worth sitting with. Standard input and output rates didn't move. What moved is the price of reading from the prompt cache — and for anyone running Claude Code against a large, mostly-static codebase context, that's the number that actually drives your bill.
Why a cache-read price cut matters more than a sticker price cut
Here's the mechanic, if you haven't dug into how Claude's prompt caching works: when a session reuses the same system prompt, file context, or tool definitions across multiple turns, Claude can cache that content instead of reprocessing it from scratch on every single call. You pay a small premium the first time content is written to the cache, then a much cheaper rate every time it's read back on a subsequent turn. For a coding agent that's repeatedly re-reading the same repo structure, style guide, or set of tool schemas across a 40-turn session, cache reads — not fresh input tokens — end up being the bulk of your token spend.
That's exactly why dropping the cache read rate from $1.00 to $0.25 per million tokens matters so much more than it looks on paper. Anthropic's own estimate is that this reduces Fable 5.1's effective cost by around 25% for typical workloads, and as much as 45% for agentic workflows that lean heavily on cached context — which describes most real Claude Code usage once you're past a trivial one-shot prompt.
The number that turns heads: cheaper cache reads than Opus 5
The detail that got the most attention in the pricing breakdown: Fable 5.1's new $0.25 cache rate is half of Opus 5's $0.50 cache rate, even though Fable's base input price is double Opus 5's. In other words, if your workload is cache-heavy — long sessions, big repeated context, an agent that reads the same files over and over — Fable 5.1 can end up cheaper per session than a "cheaper" model with a lower sticker price but a pricier cache rate. Raw per-token pricing was never the full picture for agentic work, and this release is a pretty direct demonstration of that.
The benchmark side, with the appropriate caveat
Fable 5.1 also posted real gains on agentic benchmarks: 52.6% on Terminal-Bench-Science versus 24.7% for Fable 5, 55.8% on Terminal-Bench 4.0 versus 42.0%, and 31.4% on AutomationBench versus 17.1%. Those are large jumps, and worth knowing about. They're also vendor-reported numbers, not independent third-party evals — which doesn't make them wrong, but does mean you should treat "more than double the score on Terminal-Bench-Science" as a claim to watch for independent confirmation on, not a settled fact to build a comparison post around just yet.
How it stacks up on raw pricing against the field
For context, this launched into a market where GPT-5.6 Sol runs $4 input / $20 output per million tokens and Gemini 3.7 Flash runs $0.75 / $3.75 — both meaningfully cheaper than Fable 5.1's $10 / $50 standard rate on paper. Neither of those competitors currently matches Fable's cache efficiency, though, which is the whole point: for a short, low-context task, the cheaper sticker price wins easily. For a long Claude Code session with a lot of repeated context, the cache rate is doing most of the work, and that's where Fable 5.1 quietly closes — or in some shapes of workload, reverses — the gap.
What to actually do with this if you use Claude Code
A few practical takeaways, separate from the announcement noise:
- Structure your context to maximize cache hits. Put stable content — system prompts, style guides, tool definitions, unchanging file context — early and consistently in your prompt structure so it actually gets cached and reused, rather than shuffling it around between turns in a way that invalidates the cache.
- Long sessions benefit more than short ones. If your Claude Code usage is mostly quick one-off edits, this pricing change barely touches your bill. If you're running extended multi-file refactors or long agentic debugging sessions, it's worth actually re-checking your month-over-month usage cost — the savings compound turn over turn.
- Don't assume the cheapest listed model is the cheapest in practice. A model with a lower headline rate but an expensive cache read can lose to a pricier-looking model on a cache-heavy workload. Estimate cost off your actual usage pattern, not the top-line number.
- Fewer safety interruptions means more reliable long-running agents. The reported 60% drop in Claude Code interventions is arguably as practically useful as the price cut itself if you've had sessions stall out mid-task before.
- Treat the benchmark deltas as promising, not proven. Wait for independent evals before leaning on the Terminal-Bench or AutomationBench numbers in your own decision-making.
The bigger pattern
This is the second pricing-relevant model update in as many weeks across frontier labs, and it's a useful reminder that the sticker price on a model card is increasingly not the number that determines your actual cost. Caching, context reuse, and how a provider structures incremental pricing around agentic workflows matter as much as the headline per-token rate — sometimes more. If you're budgeting for AI-assisted development at any real scale, it's worth modeling your actual session shape against the full rate card, not just the number that shows up first in the announcement.