Portal by Spotify cut my Claude Code token usage by 90%¶
Summary¶
A Spotify engineering post arguing that most of a coding agent's token spend is
not reasoning but I/O — reading files to answer a question about one method,
generating boilerplate that mirrors twenty neighbouring files — all of it fed to
an over-qualified frontier model. The fix is model routing: keep the frontier
model (Claude Code) for problems that need reasoning, and delegate the grunt work
to a cheaper worker model (Gemini 2.5 Flash in the examples). The mechanism is
two Portal by Spotify AiKA Modes — declarative
agents on an ephemeral runtime ("AWS Lambda, but for agents") — a bulk-reader
and a code-writer, plus a Claude Code plugin called
shunt that enforces the routing through
PreToolUse hooks, wrapper scripts, and skills. Benchmarked against a Java
monorepo, mean bulk-read savings were around 90%. The load-bearing idea is
not the plugin but the modes: routing becomes a configuration problem, not a
systems-engineering one.
Key takeaways¶
- Most agent I/O is not reasoning. Reading five files to answer a question about one method, or generating a test file that follows the pattern of the twenty next to it, burns thousands of tokens with near-zero reasoning — and it all goes to a frontier model that is "wildly overqualified" for it. The seat license isn't the cost; the tokens are. (Source: sources/2026-09-03-spotify-portal-by-spotify-cut-my-claude-code-token-usage-by-90)
- Cost pressure is structural, not niche. The post cites a Gartner prediction that by 2028 AI coding costs will surpass the average developer salary; a quarter of engineering leaders already spend $200–$500/dev/month on tokens, some past $2,000. Tooling pays for itself "only if you stop burning frontier tokens on work that doesn't need them" — the efficiency-frontier thesis restated for coding agents.
- AiKA Modes = declarative agents on an ephemeral runtime. A mode is a declarative agent (instructions + model + params + attached MCP tools) that runs on an ephemeral runtime — described as "AWS Lambda, but for agents": no infra to manage, no API keys, no long-running servers. Modes are callable from the Portal CLI or API, and can be public (org-wide) or private.
- Two modes do all the work.
bulk-reader(reads provided files, answers concisely, structured bullets only) andcode-writer(generates a file from a spec + reference file, matching existing patterns, code only). Both use Gemini 2.5 Flash in the examples, but themodel:field accepts any configured model. The "output only the code / no prose, no preambles" instruction matters: without it the worker wraps output in markdown fences and prose Claude then has to parse — an instance of removing unused formatting from tool output. - shunt enforces routing in three degrading layers. (1) Hooks —
PreToolUse hooks (
check-file-size,check-bash-read) fire before every tool call; aRead(orcat/head/tail/less/more) on a file over a configurable line threshold (default 350,SHUNT_MIN_LINES) is blocked and Claude is told to use the/bulk-readerskill; targeted reads and piped commands pass through. (2) Scripts —bulk-readandcode-writebash wrappers build the Portal CLI request, wrap files in XML tags, strip markdown fences, write directly to disk, and report token usage to stderr. (3) Skills — markdown files that tell Claude when/how to call the scripts. The layering degrades gracefully: even if Claude never reads the skill, the hook still blocks the expensive read (layered-graceful-degradation-enforcement). - Every delegation is one-shot and ephemeral. Nothing is stored
server-side; re-sending the same files on a follow-up question is "free where
it matters" because the corpus goes to the worker model and never enters
Claude's context. For
code-write, the generated code goes straight to disk and Claude never sees it — saving both the reference-file read and the expensive output tokens. - Modes resolve by name, own-first. A mode name resolves case-insensitively,
preferring your own mode, then your team's, then public. Fork the public
bulk-readerinto a customised version and yours automatically takes precedence — no configuration needed. - What you cannot delegate — the reasoning/I-O boundary is explicit. You can't delegate editing (worker summaries lack reliable line numbers, so Claude must read the specific section directly for edits — hence hooks allow targeted offset/limit reads). You can't delegate reasoning (the worker found surface patterns but missed a subtle thread-safety bug that Claude caught in seconds; routing explicitly excludes debugging, architectural decisions, and safety-critical code). And latency adds up — each delegation is a round-trip (Claude Code → Portal backend → worker model → back), typically 10–30s, with Portal capping a single invocation at 30s, so very large generations must be split. The line threshold exists precisely because below it, delegation overhead exceeds the savings.
Operational numbers¶
- ~90% mean token savings on the bulk-read scenario (Java monorepo, four scenarios).
- 350 lines — default
check-file-sizethreshold below which reads pass through (SHUNT_MIN_LINES, overridable via.claude/settings.json). - 10–30s typical delegation latency; 30s hard cap per Portal invocation.
- Gemini 2.5 Flash as the worker model in the published examples (model is configurable per mode).
- Cost context (cited, external): Gartner — AI coding cost > average dev salary by 2028; 25% of eng leaders at $200–$500/dev/month, some >$2,000.
Systems / concepts / patterns extracted¶
- Systems: Portal by Spotify (AiKA Modes runtime), shunt (Claude Code plugin), MCP (modes attach MCP tools), Backstage (Portal is documented in the Backstage docs tree), Claude Code as the frontier client, Gemini Flash as the worker model.
- Concepts: io-vs-reasoning-work-split, concepts/model-first-routing, concepts/model-first-routing, concepts/token-overhead, concepts/efficiency-frontier.
- Patterns: delegate-io-to-cheaper-worker-model, hook-blocks-expensive-tool-call-redirects-to-cheaper-mode, ephemeral-declarative-agent-mode, layered-graceful-degradation-enforcement, patterns/tool-surface-minimization.
Caveats¶
- Single-author, product-forward post (the title is a personal-result claim); the 90% figure is one author's benchmark on one Java monorepo across four scenarios, not an independent or fleet-wide measurement.
- The
code-writescenario is explicitly noted as hard to measure in tokens (Claude both reads references and emits output without shunt), so its saving is qualitative, not a clean percentage. - Portal by Spotify is a commercial product; the architecture (ephemeral declarative modes + client-side routing plugin) is the transferable part, not the specific vendor.
Source¶
- Original: https://engineering.atspotify.com/2026/9/portal-by-spotify-cut-my-claude-code-token-usage-by-90/
- Raw markdown:
raw/spotify/2026-09-03-portal-by-spotify-cut-my-claude-code-token-usage-by-90-7259b780.md
Related¶
- systems/portal-by-spotify
- systems/shunt-claude-code-plugin
- io-vs-reasoning-work-split
- delegate-io-to-cheaper-worker-model
- hook-blocks-expensive-tool-call-redirects-to-cheaper-mode
- ephemeral-declarative-agent-mode
- layered-graceful-degradation-enforcement
- concepts/model-first-routing
- concepts/model-first-routing
- concepts/token-overhead
- companies/spotify