How Atlassian built a scalable synthetic data engine¶
Summary¶
Atlassian's Cloud Transition organization needed to validate that its on-prem→Cloud migration tooling could handle large enterprise customers' data before those customers migrated. That required generating dozens of realistic synthetic datasets for Jira and Confluence at enterprise scale (tens of millions of Jira work items). Their existing generator drove product APIs, which are designed for real interactive users, not bulk generation — a single test environment took multiple weeks to prepare, API rate limits became a bottleneck, and adding coverage for a new entity was slow and operationally expensive. The team built Data Brewery, a config-driven synthetic data platform built on a two-phase generation model that decouples row generation from relationship (foreign-key) resolution: Phase 1 generates all non-FK fields for every table independently (and thus in parallel); Phase 2 injects FK values in dependency order using only parent rows that already exist. Relationships (simple, composite, polymorphic) are declarative config, not custom code, and large jobs are broken into chunks for bounded memory, fault isolation, and horizontal scaling. The load-bearing insight from scale testing: total data volume is not the primary bottleneck — the distribution (shape) is (one space with thousands of GB of attachments stresses a migration pipeline very differently from the same volume spread across many spaces).
Key takeaways¶
-
API-driven bulk generation does not scale. The prior tool relied on product APIs designed for interactive users; at bulk-generation scale a single test environment took multiple weeks to prepare, API rate limits became a bottleneck, and onboarding a new entity was slow and repetitive (Source: sources/2026-09-21-atlassian-how-atlassian-built-a-scalable-synthetic-data-engine).
-
A senior architect reframed the problem: don't build the whole dataset through the product — start small and extrapolate. This killed the API approach and pushed the team toward direct generation of SQL dumps from high-level, non-identifiable metadata and data shapes.
-
Traditional synthetic-data tooling assumes a single-pass model that breaks on relationally dense schemas. Single-pass = create a row, populate all fields, resolve foreign keys immediately, move on. That fails once the schema has dozens of interdependent tables, because (a) FKs are not just columns — enterprise schemas carry simple, composite, and polymorphic foreign keys, so FK assignment is about satisfying cross-table constraints, not just matching a value; and (b) generation order becomes a scalability tax — every child row must wait for a queryable parent, adding sequencing constraints, DB reads, and branching, and limiting parallelism to the entity hierarchy (you cannot parallelize space and work item generation).
-
Core idea: separate row generation from relationship resolution. During the POC, tables with no relationships generated fast; adding FKs slowed generation dramatically, and each added dependency made it worse. The fix: do not resolve foreign keys while generating the rest of the row. This is the separation-of-concerns insight applied to data generation — the expensive coupling (relational resolution) is pulled out of the throughput-sensitive hot path.
-
Data Brewery uses a two-phase generation model.
- Phase 1 — generate tables independently. Every table's rows are produced with all non-FK fields populated (IDs, timestamps, names, text, enums, derived values) per config. FK-designated columns are intentionally left unresolved. Output = a "batch without FK": structurally valid rows with no relational links yet.
-
Phase 2 — inject FKs in dependency order. A dedicated FK injector resolves and injects FK values using valid parent-side records only — rows generated earlier in the same run (held in memory) and/or rows already in the DB. Phase 2 never creates new rows; it only decides which existing parent each child points to. Output = a "batch with FK" satisfying referential integrity.
-
Requested data shape is applied in Phase 2 as percentile targets. E.g. work items per space:
min: 100 · p50: 1,000 · p90: 8,000 · p99: 50,000 · max: 500,000. Reading the distribution: every space has ≥100 work items; half have ≤1,000; 90% have ≤8,000; the top 1% exceed 50,000; the largest may have up to 500,000. Shape is a first-class, declarative knob — not an emergent side effect of row-by-row generation. -
The system is config-driven by deliberate design decision (even at the cost of significant implementation complexity), so a new entity / data shape / relationship combination needs little product-specific engineering. Key abstractions: Product Configuration (where to find table/entity/relationship declarations), Entity Definitions (a business concept spanning one or more tables, e.g. attachment / work item / comment), Table Definitions (PKs, nullability, non-FK generation behavior), Field Generators (strategies: sequential IDs, weighted enums, templated strings), and Declarative Relationships (FKs declared as simple / composite / polymorphic so the engine applies resolution logic uniformly).
-
The FK injection engine has four components. ForeignKeyRegistry loads/holds relationship definitions; ForeignKeyValueProvider fetches valid parent values from in-memory state, the DB, or both; ForeignKeyInjector orchestrates batch injection; DependencyContext stores current-run rows/IDs before persistence so a child can reference a same-run parent without a DB round-trip. Declarative relationships mean the same resolution logic applies across products and entities, relationships are visible rather than buried in custom handlers, and complex FK patterns live in one place.
-
Chunking gives bounded memory, fault isolation, and horizontal scaling. Instead of processing everything at once, Data Brewery breaks large jobs into smaller chunks: each chunk is processed and flushed independently (bounded memory — the engine never holds the full dataset), a failed chunk is retried alone (fault isolation — not the whole run), and chunks distribute across workers (horizontal scaling — throughput scales with compute). This is a micro-batching discipline applied to bulk data generation.
-
The relational chain shows FK resolution leaving the hot path. For
Users → Spaces → Work Items → Comments → Attachments, the API-driven system had to resolve space, reporter, assignee, and status at row-creation time. In Data Brewery: generate users, generate spaces independently, generate work items with all non-FK fields, then inject space/reporter/assignee references from valid rows, then generate comments and inject work-item references. Each step runs in bulk and FK injection only touches the columns that need it. -
Volume is not the bottleneck — distribution (shape) is. Scale benchmarking found that a single space with thousands of GB of attachments stresses a migration pipeline differently than the same volume spread across dozens of spaces. Testing against realistic customer data shapes (not just aggregate volume) let the team find the real bottlenecks and optimize single-space migration resilience — a data-skew argument for test-data design.
Systems / concepts / patterns extracted¶
- Data Brewery — Atlassian's config-driven synthetic data generator platform (new system page). Two-phase generation; declarative simple/composite/polymorphic FKs; chunked execution; percentile-shape targets; four-component FK injection engine. Generates SQL-dump datasets for Jira, Confluence, and extensible ecosystem-app schemas.
- Jira — one of the products whose data is generated (spaces = projects, work items = tickets, comments, attachments); the entity vocabulary Data Brewery models.
- concepts/separation-of-concerns — the core design move: decouple the expensive relational-resolution concern from the throughput-sensitive row-generation concern (two passes over one coupling).
- concepts/micro-batching — chunked processing for bounded memory + per-chunk fault isolation + horizontal scaling of a batch workload.
- concepts/partition-skew-data-skew — the "shape, not volume, is the bottleneck" lesson; percentile-distributed synthetic data is built specifically to reproduce production skew for migration testing.
- concepts/schema-evolution — surfaced as an open problem: schema drift between Data Brewery config and evolving DB schemas; mitigated with startup validation, exploring tighter schema-registry integration ("a silent mismatch is harder to catch than a compile error").
Recorded as prose/tags only (taxonomy gate — not minted as pages)¶
Per AGENTS.md taxonomy discipline, these single-source, article-specific ideas are tagged here and described in prose rather than given dedicated concept/pattern pages:
- two-phase relational data generation (generate rows without FKs, then inject FKs in dependency order) — Data Brewery's central mechanic; canonicalized into concepts/separation-of-concerns as its wiki home for now. Promote to a pattern if a 2nd source describes the same decouple-generation-from-resolution technique.
- config-driven / declarative-over-imperative generation ("push the volatile part into configuration") — recorded as prose; overlaps existing concepts/separation-of-concerns and the config-over-code principle.
- declarative simple/composite/polymorphic FK resolution — an implementation detail of Data Brewery's engine, kept on the system page.
- percentile-shape targets for test data — folded into the concepts/partition-skew-data-skew "shape drives the bottleneck" framing.
Operational numbers & specifics¶
- Prior API-driven approach: a single test environment took multiple weeks to prepare; API rate limits were the bulk-generation bottleneck.
- Target scale: tens of millions of Jira work items, concentrated in a few spaces, generating dozens of unique datasets.
- Percentile shape example (work items per space): min 100, p50 1,000, p90 8,000, p99 50,000, max 500,000.
- FK types handled: simple, composite (multi-column parent key), polymorphic (referenced table chosen by a discriminator/type field on the child).
- Datasets are emitted as SQL dumps (or similar) built from high-level, non-identifiable metadata + data shapes.
Caveats¶
- Tier-3 source (Atlassian how-we-build blog). Included because it is a genuine platform-architecture post — two-phase generation model, config-driven engine design decision, chunked horizontal scaling, and named FK-injection components with real design trade-offs.
- The article is qualitative on numbers beyond the shape example: no end-to-end generation throughput, worker counts, chunk sizes, or before/after generation-time deltas are given.
- Two open problems are explicitly unresolved: schema drift (keeping config in sync with evolving DB schemas; startup validation added, schema-registry integration still exploratory) and conditional field rules (constraints that depend on other generated values — e.g. a status field gating which sub-fields are valid — which declarative config "does not yet express cleanly").
- Component internals (ForeignKeyRegistry / ForeignKeyValueProvider / ForeignKeyInjector / DependencyContext) are named but not deeply specified (no persistence, batching-size, or memory-budget details).
Source¶
- Original: https://www.atlassian.com/blog/how-we-build/how-atlassian-built-scalable-synthetic-data-engine
- Raw markdown:
raw/atlassian/2026-09-21-how-atlassian-built-a-scalable-synthetic-data-engine-e6fe2fa0.md