Encoding Your Domain Expert: The Context Layer Behind Spotify's Data Assistant¶
Summary¶
Spotify built Vedder, a natural-language data assistant that answers
questions against a warehouse of 70,000+ datasets / petabytes / 1.4 trillion
data points per day. The core argument of the post is that the interesting
part is not the model — it's the context layer and the ownership model
that make answers trustworthy. Dumping all schemas into an LLM does not work at
this scale: context windows are finite, and schemas alone don't encode business
meaning (what "active user" means, that country holds 'US'/'GB'/'SE', or
that IDs < 100 are legacy test data). Spotify's answer is the cluster — a
data domain owned by a named team of domain experts, bundling curated datasets
(with profiling), curated question→SQL example pairs, and business docs.
Every example is expert-reviewed and marked canonical; when Spotify tried to
auto-mine question/SQL pairs from warehouse query history, curators accepted only
12.5%. Clusters carry a continuously-computed health score, and every
conversation feeds back to cluster owners — scaling one data scientist's
expertise to thousands of users.
Key takeaways¶
- Schemas don't convey meaning; a curated context layer must sit between the warehouse and the LLM. "Provide the same number of tables to a model, and it will be confident in selecting the wrong one." The gap is business semantics: definitions, gotchas, which columns to use and which to avoid. (See concepts/text-to-sql, layered-grounded-context-for-data-agent.)
- The unit of context is the cluster — an owned data domain. A cluster can be tied to an initiative, an org, or an ad-hoc interest, is owned by a named team, and consists of datasets (relevant tables with full schema + profiling: column cardinality, common-value samples, partition structure), example query pairs (canonical question→SQL), and docs (terminology, gotchas, team-varying definitions). (See cluster-model-for-domain-context.)
- Human judgment is the gate, not a shortcut around it. Auto-generating question-SQL pairs from full warehouse query history "looks like a way to scale," but curators accepted only 12.5% of proposed pairs — the other 87.5% were ad-hoc exploration, debugging sessions, one-off answers, queries using the wrong table, or "technically correct but taught the wrong pattern." "Query history is rich. Most of it is noise. And the signal doesn't label itself." (See curator-approved-example-selection, expert-curated-question-sql-pairs.)
- Profiling captures the values, not just the types. Vedder stores column
cardinality, samples of common values, and partition structure so the model
knows
countryhas values like'US','GB','SE'"rather than guessing" when it writes aWHEREclause. - Context rots; clusters need a health score. Because schemas evolve, columns get renamed, and tables get deprecated, each cluster has a continuously-computed health score built from signals: underlying-data health, fraction of curated pairs still valid after recent schema changes (a renamed column degrades referencing pairs immediately), context coverage of questions people actually ask, and SQL reproducibility. Degradation is surfaced with suggested actions on the owner's dashboard. (See context-health-score, context-coverage-as-quality-metric.)
- The agent is a ReAct loop with transparent sourcing and an "I don't know." On a question, Vedder picks context, writes SQL, runs it, and returns the answer with the query and its sources; it follows a ReAct reason/act loop adjusting on each tool call; and when no knowledge base covers the topic it says so — "That transparency is what makes the answers it gives reliable."
- Meet users where they work, multi-surface. Vedder ships as a Slack bot, an MCP server for IDEs/AI tools, and a dedicated web UI. (See systems/model-context-protocol.)
- Every conversation closes the loop. Vedder logs every conversation, query, answer, generated SQL, and user feedback and shows them to cluster owners; each approved pair / clarified doc improves the next user's answer — a data network effect over curated context. (See continuous-evaluation-feedback-loop, human-override-as-ground-truth-signal.)
- The architecture is deliberately not Spotify-specific. "The people who best understand a data domain are the best ones to curate the context the model sees." The role of the data expert shifts from answering one-off questions to shaping the knowledge layer that answers thousands.
Systems / concepts / patterns extracted¶
- System: Vedder — Spotify's NL data assistant (cluster context layer + ReAct agent + Slack/MCP/web surfaces).
- Concepts: cluster-model-for-domain-context (new), context-health-score (new), curator-approved-example-selection (new), concepts/text-to-sql, layered-grounded-context-for-data-agent, concepts/governed-agent-data-access, context-coverage-as-quality-metric, continuous-evaluation-feedback-loop, query-history-knowledge-base, concepts/human-in-the-loop, data-network-effect.
- Patterns: expert-curated-question-sql-pairs (new), human-override-as-ground-truth-signal, data-product-as-agent-context-layer.
Operational numbers¶
- Warehouse: 70,000+ datasets, petabytes, 1.4 trillion data points/day.
- Usage since August 2025: 2,100+ Spotifiers, 13,000+ conversations, 60,000+ messages, 177 clusters across advertising, podcasts, music, audiobooks, finance, creator tools, and a dozen+ other domains.
- >25% of users had never written SQL before.
- Curator acceptance of auto-mined example pairs: 12.5% accepted / 87.5% rejected.
Caveats¶
- This is a Tier-2 engineering post with a partly narrative framing; the health-score signal list is described qualitatively ("and a handful of others"), not exhaustively.
- The post is explicit that context curation is a foundation, and flags an open frontier: knowledge that lives outside the schema (process docs, org definitions) is "some of the questions we are exploring next."
- No model name, retrieval index, or SQL-dialect internals are disclosed.
Source¶
- Original: https://engineering.atspotify.com/2026/6/encoding-your-domain-expert-the-context-layer-behind-spotifys-data-assistant/
- Raw markdown:
raw/spotify/2026-06-10-encoding-your-domain-expert-the-context-layer-behind-spotify-65b0af2b.md
Related¶
- systems/spotify-vedder — the assistant described here.
- systems/databricks-genie — sibling curated-context analytical agent.
- concepts/text-to-sql — the task Vedder performs.
- cluster-model-for-domain-context — Vedder's unit of context.
- context-health-score — keeping clusters current.
- curator-approved-example-selection — the 12.5% story.
- layered-grounded-context-for-data-agent — sibling framing.
- expert-curated-question-sql-pairs — the curation pattern.
- companies/spotify