SYSTEM Cited by 1 source
Data Brewery¶
What it is¶
Data Brewery is Atlassian's synthetic data generator platform — a config-driven engine that produces realistic synthetic datasets (SQL dumps) for Jira, Confluence, and extensible ecosystem-app schemas from high-level, non-identifiable metadata and data shapes. It was built by Atlassian's Cloud Transition organization to validate that on-prem→Cloud migration tooling could handle large enterprise customers' data at scale before those customers migrated — generating dozens of unique datasets with tens of millions of Jira work items at realistic distributions (Source: sources/2026-09-21-atlassian-how-atlassian-built-a-scalable-synthetic-data-engine).
It replaced a prior generator that drove product APIs (designed for interactive users, not bulk generation), where a single test environment took multiple weeks to prepare and API rate limits were the bottleneck.
Two-phase generation model¶
Data Brewery's defining architecture decouples row generation from relationship (foreign-key) resolution — an application of separation of concerns to data generation. The observation from the POC: tables with no relationships generated fast; adding FKs slowed generation dramatically, and each added dependency made it worse. So Data Brewery does not resolve foreign keys while generating the rest of the row.
- Phase 1 — generate tables independently. Every table's rows are produced with all non-FK fields populated (IDs, timestamps, names, text, enums, derived values) per config. FK-designated columns are left unresolved. Because tables have no cross-references yet, they can be generated in parallel. Output: a batch without FK — structurally valid rows with no relational links.
- Phase 2 — inject FKs in dependency order. A dedicated FK injector resolves and injects FK values using valid parent-side records only: rows generated earlier in the same run (held in memory) and/or rows already in the DB. Phase 2 never creates new rows — it only decides which existing parent each child points to. Output: a batch with FK satisfying referential integrity.
Requested data shape is applied here as percentile targets, e.g.
work items per space: min 100 · p50 1,000 · p90 8,000 · p99 50,000 ·
max 500,000. Shape is a first-class declarative knob, not an emergent
side effect of row-by-row generation.
Why FKs are hard (the problem it solves)¶
Traditional synthetic-data tooling assumes a single-pass model (create a row, populate all fields, resolve FKs immediately, move on), which breaks on relationally dense enterprise schemas because:
- Foreign keys are not just columns. Enterprise schemas carry simple (one-column), composite (multiple columns identify the parent), and polymorphic (referenced table depends on a discriminator/type field on the child) FKs. Assignment is about satisfying cross-table constraints, not matching a value.
- Generation order becomes a scalability tax. If every row must be complete at creation, each child waits for a queryable parent — adding sequencing constraints, DB reads, and branching, and limiting parallelism to the entity hierarchy (you cannot parallelize space and work item).
Config-first design¶
A deliberate design decision (accepting significant implementation complexity) so a new entity / shape / relationship combination needs little product-specific engineering. Key abstractions:
- Product Configuration — where the engine finds table, entity, and relationship declarations.
- Entity Definitions — a business concept spanning one or more tables (e.g. attachment, work item, comment).
- Table Definitions — primary keys, nullability rules, non-FK column generation behavior.
- Field Generators — value strategies: sequential IDs, weighted enums (weighted to resemble real distributions), templated/pattern strings.
- Declarative Relationships — FKs declared as simple / composite / polymorphic so the engine applies resolution logic uniformly. Because the rules live in config, the same resolution logic applies across products and entities, relationships are visible instead of buried in custom handlers, and complex FK patterns are handled in one place.
FK injection engine¶
Four components:
- ForeignKeyRegistry — loads and holds relationship definitions.
- ForeignKeyValueProvider — fetches valid parent values from in-memory state, the database, or both.
- ForeignKeyInjector — orchestrates batch injection.
- DependencyContext — stores current-run rows and IDs before they are persisted, so a child can reference a parent generated earlier in the same execution without a database round-trip.
Scaling: chunked execution¶
Large jobs are broken into smaller chunks — a micro-batching discipline that gives:
- Bounded memory — each chunk is processed and flushed independently; the engine never holds the full dataset in memory.
- Fault isolation — a failed chunk is retried alone, not the whole run.
- Horizontal scaling — chunks distribute across workers, so throughput scales with available compute.
The relational chain (worked example)¶
Users → Spaces → Work Items → Comments → Attachments. In the old
API-driven system, generating a work item meant an endpoint that
resolved space, reporter, assignee, and status at row-creation
time. In Data Brewery: (1) generate users; (2) generate spaces
independently; (3) generate work items with all non-FK fields; (4)
inject space/reporter/assignee references from valid rows; (5) generate
comments independently, then inject work-item references. Each step runs
in bulk; FK injection only touches the columns that need it.
Load-bearing lesson: shape, not volume¶
Scale benchmarking found total data volume is not the primary bottleneck — the distribution (shape) is. One space with thousands of GB of attachments stresses a migration pipeline very differently than the same volume spread across dozens of spaces. Designing synthetic datasets against realistic customer shapes let the team find the right bottlenecks and optimize single-space migration resilience — a data-skew argument for test-data design.
Open problems¶
- Schema drift — keeping config in sync as DB schemas evolve. Startup validation catches mismatches early; tighter schema-registry integration is still being explored, because "a silent mismatch is harder to catch than a compile error."
- Conditional field rules — some entities need constraints that depend on other generated values (e.g. a status field that gates which sub-fields are valid); declarative config "does not yet express that cleanly."
Related¶
- systems/jira — a product whose data Data Brewery generates
- concepts/separation-of-concerns — decouple row generation from FK resolution
- concepts/micro-batching — chunked execution for bounded memory + scaling
- concepts/partition-skew-data-skew — shape-not-volume test-data lesson
- concepts/schema-evolution — schema-drift open problem
- companies/atlassian