Skip to content

SPOTIFY

Read original ↗

Portal by Spotify cut my Claude Code token usage by 90%

Summary

A Spotify engineering post arguing that most of a coding agent's token spend is not reasoning but I/O — reading files to answer a question about one method, generating boilerplate that mirrors twenty neighbouring files — all of it fed to an over-qualified frontier model. The fix is model routing: keep the frontier model (Claude Code) for problems that need reasoning, and delegate the grunt work to a cheaper worker model (Gemini 2.5 Flash in the examples). The mechanism is two Portal by Spotify AiKA Modes — declarative agents on an ephemeral runtime ("AWS Lambda, but for agents") — a bulk-reader and a code-writer, plus a Claude Code plugin called shunt that enforces the routing through PreToolUse hooks, wrapper scripts, and skills. Benchmarked against a Java monorepo, mean bulk-read savings were around 90%. The load-bearing idea is not the plugin but the modes: routing becomes a configuration problem, not a systems-engineering one.

Key takeaways

  • Most agent I/O is not reasoning. Reading five files to answer a question about one method, or generating a test file that follows the pattern of the twenty next to it, burns thousands of tokens with near-zero reasoning — and it all goes to a frontier model that is "wildly overqualified" for it. The seat license isn't the cost; the tokens are. (Source: sources/2026-09-03-spotify-portal-by-spotify-cut-my-claude-code-token-usage-by-90)
  • Cost pressure is structural, not niche. The post cites a Gartner prediction that by 2028 AI coding costs will surpass the average developer salary; a quarter of engineering leaders already spend $200–$500/dev/month on tokens, some past $2,000. Tooling pays for itself "only if you stop burning frontier tokens on work that doesn't need them" — the efficiency-frontier thesis restated for coding agents.
  • AiKA Modes = declarative agents on an ephemeral runtime. A mode is a declarative agent (instructions + model + params + attached MCP tools) that runs on an ephemeral runtime — described as "AWS Lambda, but for agents": no infra to manage, no API keys, no long-running servers. Modes are callable from the Portal CLI or API, and can be public (org-wide) or private.
  • Two modes do all the work. bulk-reader (reads provided files, answers concisely, structured bullets only) and code-writer (generates a file from a spec + reference file, matching existing patterns, code only). Both use Gemini 2.5 Flash in the examples, but the model: field accepts any configured model. The "output only the code / no prose, no preambles" instruction matters: without it the worker wraps output in markdown fences and prose Claude then has to parse — an instance of removing unused formatting from tool output.
  • shunt enforces routing in three degrading layers. (1) Hooks — PreToolUse hooks (check-file-size, check-bash-read) fire before every tool call; a Read (or cat/head/tail/less/more) on a file over a configurable line threshold (default 350, SHUNT_MIN_LINES) is blocked and Claude is told to use the /bulk-reader skill; targeted reads and piped commands pass through. (2) Scripts — bulk-read and code-write bash wrappers build the Portal CLI request, wrap files in XML tags, strip markdown fences, write directly to disk, and report token usage to stderr. (3) Skills — markdown files that tell Claude when/how to call the scripts. The layering degrades gracefully: even if Claude never reads the skill, the hook still blocks the expensive read (layered-graceful-degradation-enforcement).
  • Every delegation is one-shot and ephemeral. Nothing is stored server-side; re-sending the same files on a follow-up question is "free where it matters" because the corpus goes to the worker model and never enters Claude's context. For code-write, the generated code goes straight to disk and Claude never sees it — saving both the reference-file read and the expensive output tokens.
  • Modes resolve by name, own-first. A mode name resolves case-insensitively, preferring your own mode, then your team's, then public. Fork the public bulk-reader into a customised version and yours automatically takes precedence — no configuration needed.
  • What you cannot delegate — the reasoning/I-O boundary is explicit. You can't delegate editing (worker summaries lack reliable line numbers, so Claude must read the specific section directly for edits — hence hooks allow targeted offset/limit reads). You can't delegate reasoning (the worker found surface patterns but missed a subtle thread-safety bug that Claude caught in seconds; routing explicitly excludes debugging, architectural decisions, and safety-critical code). And latency adds up — each delegation is a round-trip (Claude Code → Portal backend → worker model → back), typically 10–30s, with Portal capping a single invocation at 30s, so very large generations must be split. The line threshold exists precisely because below it, delegation overhead exceeds the savings.

Operational numbers

  • ~90% mean token savings on the bulk-read scenario (Java monorepo, four scenarios).
  • 350 lines — default check-file-size threshold below which reads pass through (SHUNT_MIN_LINES, overridable via .claude/settings.json).
  • 10–30s typical delegation latency; 30s hard cap per Portal invocation.
  • Gemini 2.5 Flash as the worker model in the published examples (model is configurable per mode).
  • Cost context (cited, external): Gartner — AI coding cost > average dev salary by 2028; 25% of eng leaders at $200–$500/dev/month, some >$2,000.

Systems / concepts / patterns extracted

Caveats

  • Single-author, product-forward post (the title is a personal-result claim); the 90% figure is one author's benchmark on one Java monorepo across four scenarios, not an independent or fleet-wide measurement.
  • The code-write scenario is explicitly noted as hard to measure in tokens (Claude both reads references and emits output without shunt), so its saving is qualitative, not a clean percentage.
  • Portal by Spotify is a commercial product; the architecture (ephemeral declarative modes + client-side routing plugin) is the transferable part, not the specific vendor.

Source

Last updated · 766 distilled / 2,225 read