CONCEPT Cited by 3 sources
Text-to-SQL¶
Text-to-SQL is the task of generating an executable SQL query from a natural-language question, given a specific database schema (and, in practice, a mass of surrounding context: table docs, glossary terms, metric definitions, historical queries, freshness/quality signals).
Why the naive recipe breaks at production scale¶
The textbook shape — "give an LLM the question plus a table dump of the schema" — works for toy benchmarks but breaks on real warehouses. Pinterest names four failure modes at 100,000+ analytical tables and 2,500+ analysts scale (sources/2026-03-06-pinterest-unified-context-intent-embeddings-for-scalable-text-to-sql):
- Vocabulary mismatch. The analytical question does not match any table description's wording.
- Multiple candidate tables, one right join path. Many tables could answer the question, but only specific join patterns work.
- Company-specific metric conventions. The "right" way to compute a metric involves Pinterest-specific domain logic not visible in the schema.
- Split sources of truth. Quality signals (tiering), authoritative schemas, and established query patterns live in different systems — no single retrieval pulls all the context needed.
Two production shapes documented on the wiki¶
Shape A: unified context-intent embeddings over query history (Pinterest)¶
Pinterest's production answer (sources/2026-03-06-pinterest-unified-context-intent-embeddings-for-scalable-text-to-sql):
- Encode historical SQL queries through a three-step pipeline (domain context injection → SQL-to-text → text-to-embedding).
- Retrieve by analytical intent, not by table description.
- Rank with governance-aware fusion over tier / freshness / ownership.
- Generate SQL that reuses validated join keys + filters from the
retrieved historical patterns, then
EXPLAIN-validate beforeEXECUTE.
Thesis: "your analysts already wrote the perfect prompt." Query history is a knowledge base of expert- authored analytical solutions.
Shape B: graph-walk SQL generation (Netflix UDA / Sphere)¶
Netflix's contrasting answer (sources/2025-06-14-netflix-model-once-represent-everywhere-uda; graph-walk-sql-generation):
- Model business concepts + data containers as nodes in a knowledge graph; mappings are edges.
- User picks concept endpoints; graph-walk produces the JOIN specification.
- SQL is compiled from the walk, not generated by an LLM.
The two shapes are not competitors so much as different stances on where the domain knowledge lives: Pinterest treats query history as the knowledge substrate; Netflix treats the curated knowledge graph as the substrate.
What production Text-to-SQL needs beyond "LLM + schema"¶
Pinterest explicitly names the infrastructure surfaces required:
- Governance + tiering to separate production-grade tables from staging/legacy.
- Glossary terms + lineage- based propagation for column-level semantic unification.
- Rich documentation — maintained at scale via AI-generated docs with a human-review ladder.
- Query-history indexing — turn accumulated analyst queries into a searchable library.
- Vector DB platform — an internal substrate so every LLM feature isn't reinventing its own index.
- Validation —
EXPLAIN-before-EXECUTE, bounded retry, conservativeLIMIT. - Asset-first design — surface existing curated assets before generating new SQL.
Seen in¶
- sources/2026-09-25-databricks-from-data-to-dialogue-how-sp-global-energy-made-its-structured-data-estate-conversational — S&P Global Energy: SME-curated semantic layer as the production shape. A Genie Agent does text-to-SQL over one dataset group, but the differentiator is that domain experts (not engineers) enrich each agent with table/column descriptions, example queries, trusted assets for high-stakes metrics, and business definitions — e.g. "floating storage is defined as cargoes idling for 3 days or more in vessels travelling below a threshold speed." "This is the step that generic text-to-SQL solutions skip — and it is the step that determines whether users trust the answers." Echoes Spotify Vedder's expert-vetted substrate finding (curation, not raw query history) and Pinterest's company-specific-metric failure mode. Best leading indicator of adoption they tracked: how often SMEs agreed with Genie's generated SQL during curation — "Measure trust, not just latency." The curated agents are then exposed as governed MCP servers (see systems/model-context-protocol).
- sources/2026-07-28-wix-how-a-tool-i-built-for-my-team-became-wixs-official-bi-platform — Wix Vizion: an agent writes SQL from a plain-language request, grounded by MCP context (so it understands what a table means "instead of guessing"), then runs it on live data via Trino. Trust is enforced by a CI check that the query actually executes and returns a sane result (see ci-gated-dashboard-publish).
- sources/2026-03-06-pinterest-unified-context-intent-embeddings-for-scalable-text-to-sql — canonical wiki introduction (Pinterest Analytics Agent).
- sources/2025-06-14-netflix-model-once-represent-everywhere-uda — the graph-walk alternative (Netflix Sphere).
- sources/2026-06-10-spotify-encoding-your-domain-expert-the-context-layer-behind-spotify-65b0af2b — Spotify Vedder: a fourth production shape where the substrate is the team-owned cluster — bundled datasets+profiling, expert-curated question→SQL pairs, and docs, kept fresh by a health score. Its distinctive finding is that raw query history is mostly noise (12.5% curator acceptance), so the domain knowledge lives in expert-vetted examples, not the full log.
Related¶
- systems/pinterest-analytics-agent
- sql-to-intent-encoding-pipeline
- retrieve-then-rank-llm
- graph-walk-sql-generation
Shape C: semantic template cache over query structure (AWS)¶
AWS documents a third production shape that separates query reuse from query generation. Its Text2SQL template cache stores a parameterized SQL structure with the embedding of the question that created it. At request time, semantic retrieval finds a matching structure; validated named entities fill typed slots through prepared parameters and the query executes against live data. The fast path therefore skips large-context SQL generation without caching a stale answer.
A low-similarity match, invalid slot, or result that fails a sufficiency check takes the authoritative full-generation path. A successful generated query is generalized and reinserted, making coverage a function of actual query demand. This is distinct from Pinterest's historical-query retrieval: Pinterest returns validated examples to inform SQL generation, while AWS can execute the selected vetted template directly. (Source: sources/2026-08-13-aws-reducing-text2sql-latency-with-parameterized-query-templates.)
See query-structure-caching, patterns/cheap-approximator-with-expensive-fallback, and parameterized-template-execution-with-entity-validation.
Merged aliases¶
sql-to-text-transformation