Skip to content

SYSTEM Cited by 1 source

Vedder (Spotify data assistant)

Vedder is Spotify's internal natural-language data assistant: ask a question in plain English and get reliable data back within seconds, along with the SQL query it ran and the sources it used. It exists to close the gap that opened as data demand outran the supply of human data experts — with 70,000+ datasets and petabytes of data, "no single individual can claim knowledge of everything," and "just putting all schemas into an LLM doesn't work at this scale." (Source: sources/2026-06-10-spotify-encoding-your-domain-expert-the-context-layer-behind-spotify-65b0af2b)

Vedder is a canonical wiki instance of a text-to-SQL data agent whose defining bet is on a curated, owned context layer rather than on the model.

What it does

When a question comes in, the agent:

  1. Picks the appropriate context (the relevant cluster — see below).
  2. Writes the SQL query.
  3. Runs it against the warehouse.
  4. Returns the answer alongside the query and its sources.

It follows a ReAct loop — reasoning and acting in steps, adjusting based on what each tool call returns — so a user can read how a result was produced, not just what it was. When no knowledge base covers the topic, Vedder says so rather than guessing; per the post, "that transparency is what makes the answers it gives reliable."

The cluster model (its context layer)

Spotify calls its data domains clusters. A cluster is owned by a named team of domain experts and can be tied to an initiative, an org, or an ad-hoc interest; the platform tells a team if a domain is already covered. Each cluster bundles three components — the general shape canonicalised at cluster-model-for-domain-context:

  • Datasets — the relevant warehouse tables with full schema and profiling: column cardinality, samples of common values, and partition structure. Profiling is load-bearing — when the model writes a WHERE clause "it helps to know that country has values like 'US', 'GB', 'SE' rather than guessing."
  • Example query pairs — canonical question→SQL examples, each reviewed and marked canonical by the cluster curators (see below).
  • Docs — additional business context: terminology, gotchas, definitions that vary by team, which columns to use and which to avoid.

Curation is owned by the data scientists and analytics engineers who know how the data is modeled — they decide how to split a domain into clusters, which tables to include, and which examples matter. This is the data-product-as-agent-context shape with an explicit ownership contract.

Human judgment over auto-mined examples

Spotify tried the obvious shortcut: the warehouse holds the complete query history of every data expert, so you can take each query, ask an LLM to infer the question it answered, and use those pairs to teach SQL generation. When curators were shown real question/SQL pairs mined from their domain, they accepted only 12.5%. The other 87.5% were ad-hoc exploration, debugging sessions, one-off answers, queries against the wrong table, or "technically correct but taught the wrong pattern." "Query history is rich. Most of it is noise. And the signal doesn't label itself." Every example therefore runs through an expert — the model reasons over context, but the experts decide what is true about the data. See curator-approved-example-selection and expert-curated-question-sql-pairs.

Keeping clusters healthy

Because "context that was accurate last month can be wrong today," each cluster carries a continuously-computed health score built from monitored signals: health of the underlying data; how many curated pairs are still valid after recent schema changes (a renamed column degrades referencing pairs immediately); how well context covers the questions people actually ask; SQL reproducibility; "and a handful of others." When any degrades, the score reflects it and actions are suggested on the owner's cluster dashboard so experts know where to spend curation time. See context-health-score.

Surfaces

Vedder is built into the surfaces people already use:

  • a Slack bot for quick questions in a thread,
  • an MCP server for IDEs and AI tools, and
  • a dedicated web UI for interactive exploration.

Closing the loop

Vedder logs every conversation, query, answer, generated SQL, and user feedback, and surfaces them to cluster owners. Every approved question-SQL pair and every clarified doc makes the next user's answer more accurate — a data network effect over curated context, and a continuous feedback loop in which human corrections become the ground-truth signal. The data expert's role shifts from answering one-off questions to shaping the knowledge layer that answers thousands.

Operational numbers

  • Warehouse: 70,000+ datasets, petabytes, 1.4 trillion data points/day.
  • Since August 2025: 2,100+ users, 13,000+ conversations, 60,000+ messages, 177 clusters.
  • >25% of users had never written SQL.

Sibling systems on the wiki

  • Databricks Genie — curated semantic context + verified example queries; the closest analogue (owner-curated example queries
  • trusted-asset scoping).
  • Cloudflare Skipper — the five-layer layered grounded context recipe; Vedder's cluster is a team-owned bundling of similar layers.
  • Pinterest Analytics Agent — treats historical query logs as the substrate; Vedder's 12.5% finding is the counterpoint on how much of that history is usable unfiltered.

Source

Last updated · 766 distilled / 2,225 read