Bryn Flow
Deep Dive

How Spotify Cut Claude Code Token Usage by 90% With a Two-Model Router

RS
Raiyan Shahid Building ExamAI & FileForge under Bryn Flow · Get in touch

A Spotify engineering post climbed into the top few stories on Hacker News this week for a claim that sounds almost too good to check: a 90% cut in Claude Code token usage, achieved with a plugin instead of a smaller model or a shorter prompt. The number holds up once you see the mechanism behind it, and the underlying pattern — split expensive reasoning from cheap I/O, and route each to the model built for it — is worth understanding whether or not you ever touch Spotify's code.

The real cost problem: it's I/O, not reasoning

The engineer behind the post frames the issue bluntly: most of what an AI coding agent does isn't thinking, it's I/O — reading a handful of large files to answer one question, generating boilerplate that already follows an established pattern, updating documentation that mirrors code that already exists. None of that needs a frontier model's reasoning, but all of it gets billed at frontier-model prices anyway, because by default every file an agent reads lands in the same context window as the reasoning it does with that file. Spotify's post cites engineering leaders already spending $2,000+ a month per developer on agent tokens, with industry projections putting AI coding costs ahead of average developer salaries by 2028 if usage patterns don't change. That's not a rounding error — it's a line item.

What Spotify actually built

The fix is a Claude Code plugin called shunt, part of a broader internal platform Spotify calls Portal, built on AiKA Modes — declarative agents that run on ephemeral infrastructure, so using one doesn't mean standing up a server or managing API keys. Two modes do the heavy lifting:

Picture a typical task: understanding how error handling is implemented across a 40-file Java module before writing a new endpoint that needs to follow the same pattern. Read directly, that's dozens of files landing in Claude's context, most of them read once and mostly discarded once the pattern is understood. Routed through bulk-reader, Gemini 2.5 Flash reads all 40 files and returns a structured summary of the pattern — a few hundred tokens instead of the full source of every file — and Claude only spends its own context on the one or two files it actually needs to edit. The heavy, disposable reading happens on a cheaper model; the precise editing happens on the model that's actually good at it.

The routing between "let Claude handle this" and "shunt this to a cheaper model" happens in three layers. A pre-tool-use hook intercepts any file read past a configurable line threshold (350 lines by default) and redirects it to bulk-reader instead of letting Claude read it directly. Bash wrapper scripts (bulk-read, code-write) handle the actual Portal CLI calls and parse the responses back into something Claude can use. And a set of markdown-based skills document the delegation syntax so Claude knows when and how to hand work off on its own. The shunt plugin sits on top of all three layers and enforces the routing decision — an expensive read never reaches Claude's context in the first place, rather than reaching it and then being summarized after the fact.

The results, and the honest limitations

Tested against a Java monorepo, bulk-read operations saw the headline 90% token reduction, and code-writer scenarios saw comparable savings since generated boilerplate skips Claude's context entirely. The tradeoff is latency: each delegated call adds 10–30 seconds, capped at 30 seconds total, since a second model has to actually run. Spotify's own writeup is upfront about where the approach doesn't apply — it can't delegate actual code edits, because a separate model can't reliably hand back the exact line numbers Claude needs to make a precise change; it can't delegate complex reasoning or debugging, which is the part of the job that actually needs a frontier model; and for small reads, the network round-trip to a second model costs more than it saves. This is a tool for bulk context-gathering and boilerplate, not a blanket replacement for using Claude directly.

Why this matters beyond one company's plugin

Both modes and the shunt plugin are public on Spotify's portal-ai-plugins repo and usable without customization, which is worth trying if you're already deep in the Claude Code ecosystem. But the more durable takeaway is the pattern itself, not the specific plugin: routing cheap, high-volume I/O to a fast model and reserving your frontier model for the reasoning it's actually good at is becoming a standard shape for agentic coding tools generally, not just Claude Code. If you use Cursor, Copilot, or anything built on the same class of coding agent, the same question is worth asking about your own workflow — how much of what you're paying frontier-model prices for is actually reasoning, and how much is just the agent reading things?

If you want to try Spotify's version directly: shunt and both AiKA modes are listed in Spotify's public plugin marketplace, installable into an existing Claude Code setup without writing custom infrastructure — though as with any third-party plugin, it's worth reading what it actually does to your tool calls before pointing it at a production codebase.

What you can do this week, without installing anything

You don't need Spotify's specific plugin to apply the same thinking to your own setup:

None of this requires you to trust a 90% number blindly — it's a reason to go measure your own agent's read-versus-write token split before assuming your workflow is already efficient.

Where to go next

Ready to start building?

Whichever language you pick, the Learn Hub has full project tutorials, cheat sheets, and interview prep to back it up.

Explore the Learn Hub