A Spotify engineering post climbed into the top few stories on Hacker News this week for a claim that sounds almost too good to check: a 90% cut in Claude Code token usage, achieved with a plugin instead of a smaller model or a shorter prompt. The number holds up once you see the mechanism behind it, and the underlying pattern — split expensive reasoning from cheap I/O, and route each to the model built for it — is worth understanding whether or not you ever touch Spotify's code.
The real cost problem: it's I/O, not reasoning
The engineer behind the post frames the issue bluntly: most of what an AI coding agent does isn't thinking, it's I/O — reading a handful of large files to answer one question, generating boilerplate that already follows an established pattern, updating documentation that mirrors code that already exists. None of that needs a frontier model's reasoning, but all of it gets billed at frontier-model prices anyway, because by default every file an agent reads lands in the same context window as the reasoning it does with that file. Spotify's post cites engineering leaders already spending $2,000+ a month per developer on agent tokens, with industry projections putting AI coding costs ahead of average developer salaries by 2028 if usage patterns don't change. That's not a rounding error — it's a line item.
What Spotify actually built
The fix is a Claude Code plugin called shunt, part of a broader internal platform Spotify calls Portal, built on AiKA Modes — declarative agents that run on ephemeral infrastructure, so using one doesn't mean standing up a server or managing API keys. Two modes do the heavy lifting:
- bulk-reader — hands multiple large files to Gemini 2.5 Flash, which returns a structured analysis instead of the raw file contents, so the files themselves never enter Claude's context window.
- code-writer — generates boilerplate that matches the surrounding codebase's existing patterns and writes it straight to disk, again without Claude ever reading or producing the file content itself.
Picture a typical task: understanding how error handling is implemented across a 40-file Java module before writing a new endpoint that needs to follow the same pattern. Read directly, that's dozens of files landing in Claude's context, most of them read once and mostly discarded once the pattern is understood. Routed through bulk-reader, Gemini 2.5 Flash reads all 40 files and returns a structured summary of the pattern — a few hundred tokens instead of the full source of every file — and Claude only spends its own context on the one or two files it actually needs to edit. The heavy, disposable reading happens on a cheaper model; the precise editing happens on the model that's actually good at it.
The routing between "let Claude handle this" and "shunt this to a cheaper model" happens in three layers. A pre-tool-use hook intercepts any file read past a configurable line threshold (350 lines by default) and redirects it to bulk-reader instead of letting Claude read it directly. Bash wrapper scripts (bulk-read, code-write) handle the actual Portal CLI calls and parse the responses back into something Claude can use. And a set of markdown-based skills document the delegation syntax so Claude knows when and how to hand work off on its own. The shunt plugin sits on top of all three layers and enforces the routing decision — an expensive read never reaches Claude's context in the first place, rather than reaching it and then being summarized after the fact.
The results, and the honest limitations
Tested against a Java monorepo, bulk-read operations saw the headline 90% token reduction, and code-writer scenarios saw comparable savings since generated boilerplate skips Claude's context entirely. The tradeoff is latency: each delegated call adds 10–30 seconds, capped at 30 seconds total, since a second model has to actually run. Spotify's own writeup is upfront about where the approach doesn't apply — it can't delegate actual code edits, because a separate model can't reliably hand back the exact line numbers Claude needs to make a precise change; it can't delegate complex reasoning or debugging, which is the part of the job that actually needs a frontier model; and for small reads, the network round-trip to a second model costs more than it saves. This is a tool for bulk context-gathering and boilerplate, not a blanket replacement for using Claude directly.
Why this matters beyond one company's plugin
Both modes and the shunt plugin are public on Spotify's portal-ai-plugins repo and usable without customization, which is worth trying if you're already deep in the Claude Code ecosystem. But the more durable takeaway is the pattern itself, not the specific plugin: routing cheap, high-volume I/O to a fast model and reserving your frontier model for the reasoning it's actually good at is becoming a standard shape for agentic coding tools generally, not just Claude Code. If you use Cursor, Copilot, or anything built on the same class of coding agent, the same question is worth asking about your own workflow — how much of what you're paying frontier-model prices for is actually reasoning, and how much is just the agent reading things?
If you want to try Spotify's version directly: shunt and both AiKA modes are listed in Spotify's public plugin marketplace, installable into an existing Claude Code setup without writing custom infrastructure — though as with any third-party plugin, it's worth reading what it actually does to your tool calls before pointing it at a production codebase.
What you can do this week, without installing anything
You don't need Spotify's specific plugin to apply the same thinking to your own setup:
- Scope what the agent can see. Exclude generated code, build output, and vendored dependencies from what your agent scans by default — every file it reads "just in case" is billed the same as a file it actually needed.
- Search before you read. Ask the agent to grep or search for the specific symbol or pattern first, and only read the surrounding lines — not the whole file — once it knows where to look.
- Use hooks to cap read size. Claude Code's hooks can intercept a tool call before it runs; a simple pre-tool-use hook that flags or blocks reads past a line-count threshold is most of what Spotify's routing layer does, without needing a second model behind it.
- Separate exploration from editing. A task that means "read a dozen files to understand how this system works" and a task that means "write this one function" have very different token profiles — treat them as separate steps rather than one long session that reads everything up front.
- Watch your own usage. Most coding agents expose token or cost usage per session somewhere. If a session is burning tokens on reads rather than edits, that's the exact signal Spotify's whole plugin was built to act on automatically.
None of this requires you to trust a 90% number blindly — it's a reason to go measure your own agent's read-versus-write token split before assuming your workflow is already efficient.