---
title: How Atlassian built a scalable synthetic data engine
source: Atlassian Engineering
source_slug: atlassian
url: https://www.atlassian.com/blog/how-we-build/how-atlassian-built-scalable-synthetic-data-engine
published: 2026-09-21
fetched: 2026-09-21T14:07:36+00:00
ingested: true
---

**Term**| **Meaning**  
---|---  
Product| An application whose data we generate eg. Jira, Confluence etc.  
Entity| Product level business object/data type (eg. _user, space, work item, comment, page_).  
Space| Jira container for work (formerly Jira  _project_)  
Work Item| Jira ticket(a bug, story, epic or task). Work items live in a space  
Attachment| A file on a Jira work item (a screenshot, log etc.)  
Data shape| Shape of data as described by sizes and constraints e.g. {count_entity_1 = 5, min_character_length_entity_1_content=100, max_entity_1_count_per_entity_2 = 9,000 etc.}  
Dataset| A set of ‘SQL dumps’ (or similar) that represents customer’s data of a given Atlassian product. Built from high-level, non-identifiable metadata and data shapes.  
Data Brewery| Atlassian’s synthetic data generator platform  
  
In this article, we describe our journey building a data-generation platform for producing synthetic data for Jira, Confluence and extensible schemas across ecosystem apps. We work in Atlassian’s Cloud Transition organization, where one of our major goal is migrating Atlassian products and ecosystem apps from on-premises to Cloud. At the time, many large enterprise customers were preparing to migrate, and we needed confidence that our migration tooling could handle their data at scale.

Before customers began their migration journeys, we wanted to answer an important question:

> _**How do we ensure their migration is fast, complete and correct?**_

To answer it, we analyzed data shapes from several upcoming customers, generated a representative dataset, and ran migration tests against it. It worked, but one shape was not enough: tens of millions of Jira _work items_ in a few _spaces_ stress a system differently from the same volume spread across many spaces.

As the number of enterprise customers preparing for migration grew, we needed data generation that could keep pace, producing dozens of unique datasets.

## Why our existing system couldn’t scale

Our existing synthetic data generation tooling relied heavily on product APIs, this model worked well enough for smaller or narrower data shapes, but these APIs were designed for real users working at scale, not for bulk-generating years of data.

As we tried generating larger Jira and Confluence datasets, several patterns became impossible to ignore:

  * Preparing a single test environment could take multiple weeks
  * API rate limits rightly exists to protect the system, but they became a bottleneck for this bulk-generation use case.
  * Adding coverage for a new _entity_ (_work item_ , _comment_ , or _attachment_) was slow and operationally expensive.



## Finding a better way to generate data

We discussed this problem with a senior architect on the team. His advice was simple:

> Stop trying to create the entire dataset through the product. Instead, start with a small dataset and figure out how to extrapolate it.

By this point, we knew the API-based approach would not work for bulk generation. It was too slow, and too dependent on product behavior that was never designed for this kind of load.

We began searching for solutions available on the market. After some research, one tool looked promising and worked well for our initial POC. But as we added more relationships, the generation time kept growing, and it was clear this would not work for the scale we needed.

## Why traditional synthetic data generation breaks on relational systems

Generating synthetic data for a single table is easy. Generating it for dozens of interdependent tables is not. Most synthetic data tooling assumes some version of a **single-pass model** : create a row, populate all fields, resolve foreign keys immediately, and move on. That model starts to fail once the schema becomes large and relationally dense.

There are a few reasons for that.

### 1\. Foreign keys are not just columns

In enterprise schemas, foreign keys often carry structural meaning. Some are simple one-column references. Others are composite foreign keys, where multiple columns together identify the parent row. Others are polymorphic foreign keys, where the referenced table depends on a discriminator or type field on the child row.

Foreign key assignment is more than finding a matching value. It is about satisfying constraints across tables.

![](https://atlassianblog.wpengine.com/wp-content/uploads/2026/09/screenshot-2026-09-21-at-2.23.59-pm.png)FK types

### 2\. Generation order becomes a scalability tax

If every row must be complete the moment it is created, then each child row must wait for parent-side values to exist and be queryable. That introduces sequencing constraints, more database reads, more branching logic, and less opportunity for independent table generation. Parallelism is limited by entity hierarchy. For example, we can’t parallelize _space_ and _work item_.

## The core idea: separate row generation from relationship resolution

While building the POC, we noticed something interesting. When tables had no relationships between them, generation was fast. The moment we introduced foreign keys between tables, things slowed down significantly. The more dependencies we added, the worse it got. This inspired a key idea:

> **Do not resolve foreign keys while generating the rest of the row.**

This led to Data Brewery, which uses a **two-phase generation model**.

![](https://atlassianblog.wpengine.com/wp-content/uploads/2026/09/image-20260318-091627.webp)Two-phase generation model

### Phase 1: Generate tables independently

In the first phase, Data Brewery generates rows for each table with all non-foreign-key fields populated. IDs, timestamps, names, text fields, enums, and derived values are generated according to configuration. Foreign-key-designated columns are intentionally left unresolved.

The output of this phase is a **batch without FK** : structurally valid rows whose relational links have not yet been injected.

### Phase 2: Inject foreign keys in dependency order

In the second phase, a dedicated foreign key injector resolves and injects FK values using valid parent-side records. These can come from:

  * Rows generated earlier in the same run and held in memory.
  * Rows already present in the database.
  * A combination of both, depending on the relationship definition.



Phase 2 never creates new rows. Every parent it links to already exists, either from Phase 1 or from the database. All it decides is which parent each child row points to and that is exactly where the requested data shape is applied. The shape is expressed as percentile targets.
    
    
    work items per space distribution
    min: 100   p50: 1,000   p90: 8,000   p99: 50,000   max: 500,000

**Distribution**| **Value**| **Meaning**  
---|---|---  
Min| 100 work items| Every _space_ has at least 100 _work items_.  
p50 / Median| 1,000 work items| Half of _spaces_ have 1,000 _work items_ or fewer.  
p90| 8,000 work items| 90% of _spaces_ have 8,000 _work items_ or fewer.  
p99| 50,000 work items| The top 1% of _spaces_ exceed 50,000 _work items_.  
Max| 500,000 work items| The largest _spaces_ may have at most 500,000 _work item_ s.  
  
The output is a **batch with FK** : rows that now satisfy referential integrity.

That separation changes the performance profile of the whole system.

Instead of resolving relationships every time it creates a row, Data Brewery first generates content quickly, then checks and fixes relationships in a separate step. This reduces work for each row, improves batching, and makes complex relationships easier to define.

## Architecture overview

![](https://atlassianblog.wpengine.com/wp-content/uploads/2026/09/screenshot-2026-09-21-at-2.26.53-pm.png)High Level Architecture

## Designing the system as configuration first, not code first

Adding a new entity support in the existing API-based system was monotonous and time-consuming since we had to integrate new APIs, the same tedious process, repeated over and over.

So from day one, the goal was not just to generate data faster, but also to reduce the amount of product-specific engineering required whenever a team needed a new entity, a new data shape, or a new combination of relationships.

This led us to one of our biggest design decisions: making Data Brewery **config-driven** , even if it added significant implementation complexity.

Below are the key abstractions:

  * **Product Configuration:** Defines where the engine finds table, entity, and relationship declarations.
  * **Entity Definitions:** Represents a business concept spanning one or more tables (e.g., _attachment_ , _work item,_ _comment_).
  * **Table Definitions:** Describes primary keys, nullability rules, and generation behavior for non-FK columns.
  * **Field Generators:** Strategies for producing values like sequential IDs, weighted enums, or templated strings.
  * **Declarative Relationships:** Defines FKs as simple, composite, or polymorphic, allowing the engine to apply resolution logic uniformly.



* * *

#### A generic domain model for generating data for any domain.

![](https://atlassianblog.wpengine.com/wp-content/uploads/2026/09/screenshot-2026-09-21-at-2.28.17-pm.png)
    
    
    name: <entity-name>
    fields:
      <primary-key-field>:
        generation: { generatorType: sequential }
      <text-field>:
        generation: { generatorType: pattern, pattern: "<Prefix>-{rowIndex}" }
      <enum-field>:
        generation:
          generatorType: value
          values:  ["<state-a>", "<state-b>", "<state-c>", "<state-d>"]
          weights: [0.3, 0.3, 0.2, 0.2]        # weighted to resemble real distributions
      <foreign-key-field>:
        schema: { columnTypes: [FOREIGN_KEY] }  # resolved later by the FK injector

## How Data Brewery keeps synthetic data linked correctly

Separating foreign-key injection from row generation only works if injection still produces valid links. That includes simple FKs, composite FKs, and polymorphic FKs.

FK injection engine resolves those parent values and writes them in batches.

Injection engine relies on four main components:

![](https://atlassianblog.wpengine.com/wp-content/uploads/2026/09/cb76281f-f129-4358-87fe-e5a4ffe41310.webp)

  * **ForeignKeyRegistry** to load and hold relationship definitions.
  * **ForeignKeyValueProvider** fetches valid parent values from in-memory state, the database, or both.
  * **ForeignKeyInjector** orchestrate batch injection.
  * **DependencyContext** stores rows and IDs from the current run before they are persisted, so a child can reference a parent generated earlier in the same execution without a database round-trip.



### Why declarative relationships matter

When those rules live in configuration, the same resolution logic applies across products and entities. Relationships are visible instead of buried in custom handlers, new entities are easier to onboard, and complex FK patterns are handled in one place.

## Scaling to millions of rows without turning the generator into a memory problem

Relational correctness was only half the problem. The other half was scale.

Instead of processing everything at once, Data Brewery breaks large jobs into smaller chunks.

  * **Bounded memory usage:** each chunk is processed and flushed independently, so the engine never holds the full dataset in memory.
  * **Fault isolation:** if a chunk fails, only that chunk is retried, not the entire run.
  * **Horizontal scaling:** chunks can be distributed across workers, so throughput scales with available compute.

![](https://atlassianblog.wpengine.com/wp-content/uploads/2026/09/a1846e8d-3a19-4da2-8e12-5224febf4997.webp)

## The relational chain

Consider a simplified dependency chain: `Users → Spaces → Work Items → Comments → Attachments`

In our previous API-driven system, generating a _work item_ meant calling an endpoint that had to simultaneously resolve the _space_ , _reporter_ , _assignee_ , and _status_ behavior, all at row creation time.

In Data Brewery, the sequence removes foreign key resolution from the throughput-sensitive hot path entirely:

  1. Generate _user_ rows with IDs and attributes.
  2. Generate _space_ rows independently.
  3. Generate _work item_ rows with all non-FK fields populated.
  4. Inject _space_ , _reporter_ , and _assignee_ relationships into _work item_ using valid referenced rows.
  5. Generate _comment_ rows independently, then inject _work item_ references.



Each step runs in bulk, and FK injection only touches the columns that need it.

## **Open engineering questions we’re still solving**

No meaningful platform story is complete without the trade-offs. Moving complexity from code into configuration does not eliminate it; it relocates the problem.

A couple of areas remain important in the platform’s evolution:

  * **Schema drift:** keeping config in sync as database schemas evolve, we added startup validation to catch mismatches early, but still exploring tighter integration with schema registries; as a silent mismatch is harder to catch than a compile error.
  * **Conditional field rules:** some entities need constraints that depend on other generated values (e.g. a status field that gates which sub-fields are valid), and declarative config does not yet express that cleanly.



These are good problems to have. They show that the architecture has moved past the “can we do this at all?” stage and into the more interesting platform question: how do we make this easy and repeatable for many teams?

## Lessons that **generalize**

Several of these stuck with us well beyond this project:

  * **Design the system for actual problem.** Like in _data-brewery’s_ case the problem was never just generating data. It was to preserve relationships at scale.
  * **Push the volatile part into configuration.** When the _what_ changes far more often than the _how_ , configuration works better than code, and turns bottleneck into a self-serve capability.
  * **Decouple the expensive coupling.** How separating row generation from relationship resolution improved the overall data generation time.
  * **Testing against close to real customer data shape.** During testing and scale benchmarking, we discovered that total data volume alone isn’t the primary bottleneck, it’s the distribution. A single space with thousands of gigabytes in attachments stresses migration pipelines differently than the same volume distributed across dozens of spaces. Designing synthetic datasets allowed us to identify the right bottlenecks and optimize single-space migration resilience.
