Skip to content

Atlassian

Atlassian Engineering blog (atlassian.com/blog/how-we-build). Tier-3 source on the sysdesign-wiki. Much of the blog is product-marketing / feature-announcement content for Jira / Confluence / Bitbucket / Rovo — which is filtered out on ingest (see wiki/log.md). When the blog does go architectural, four axes are in scope:

  • Confluence's web-tier performance — React-18 streaming SSR, hydration semantics, edge-proxy buffer-kill; 2026-04-16 streaming-SSR post is the canonical wiki instance.
  • Rovo Chat's agent orchestration architecture (Long Horizon) — single-LLM iterative reasoning loop replacing a hierarchical multi-agent orchestrator; flattened tool surface with progressive disclosure via meta-tools; context compaction service; prompt layer ordering for prefix-cache maximisation; child instances for parallel research; adaptive reasoning effort. The 2026-06-18 "Long Horizon" post is the canonical wiki instance.
  • Rovo's agent harness / distributed execution runtime — the Rovo Agent Harness is the runtime beneath Long Horizon that turns Rovo Chat into an always-on sidekick. A split-plane architecture (durable control plane serving millions of users + disposable on-demand compute sandbox); programmatic tool calling / "code mode" in the sandbox (>50% latency, 55% token, +30% accuracy vs a meta-tool baseline on complex Jira queries); dynamic model selection (skip the sandbox for lightweight queries); a callback bridge routing all sandbox tool calls back through the control plane for auth (no direct sandbox egress); a CI-versioned localized tool package on sub-100ms warm start
  • runtime active manifest; durable interactive sub-agents (a sub-agent is a hidden Rovo Chat conversation); discovery-probe + checkpoint error recovery; and a human-approved self-evolution loop. The 2026-08-27 "Opening the Door to Agent Autonomy" post is the canonical wiki instance.
  • Atlassian's ML platform (ML Studio) — unified enterprise-scale ML development platform with composable versioned modules, workflow orchestrator (hot clusters, deterministic caching, nested workflows), and embedded multi-layer compliance (column-level data classification with automatic tag propagation); serves ~120k monthly workflow runs across 100+ ML teams and is the backbone for Rovo Search/Chat, Teamwork Graph, and Confluence AI. The 2026-06-10 ML-platform-architecture post is the canonical wiki instance.
  • Atlassian's agentic-AI platform / Rovo Dev ecosystem — Fireworks (the Firecracker-microVM-on-Kubernetes secure execution substrate for AI agents), the Rovo Dev agent itself, Bitbucket Pipelines as automated quality gate for agent PRs, and the agentic-development workflow Atlassian runs internally (dev-shard iteration, AI-written e2e tests, adversarial-review sub-agents, orchestration meta-skills). The 2026-04-24 Rovo-Dev-Driven-Development post is the canonical wiki instance.
  • Search Platform / OpenSearch infrastructure — the ~150-person Search Platform org managing 300+ OpenSearch clusters, 2500+ data nodes, 1.7 PB of data across 13 regions. Powers Quick Search, Full Page Search, Rovo Search, Smart Answers, and Rovo Chat. Built AOSC for zero-downtime index reshaping. The 2026-07-24 AOSC post is the canonical wiki instance.
  • Event streaming at extreme scale (StreamHub) — StreamHub is Atlassian's central event streaming platform, handling 150B events/day (3.2M events/sec peak) via Kafka on MSK with tiered storage. Migrated from Kinesis; hardened with cluster sharding, failover-cluster escape hatches, companion-region DR, ingress rate limiting, quarantine, and client quotas. The 2026-07-28 StreamHub post is the canonical wiki instance.
  • Multi-tenant event delivery at scale (Events Rail) — Events Rail (webhooks-processor) is the platform that evaluates, matches, and delivers 10B+ events/month to Forge apps, customer webhooks, and audit-log endpoints. Two-stage pipeline (matching → delivery) with layered fairness: negative-lookup cache, queue isolation per pipeline, per-tenant rate limiting, per-recipient concurrency limiting. The 2026-07-29 Events Rail post is the canonical wiki instance.
  • Synthetic data generation for migration testing (Data Brewery) — Data Brewery is Atlassian Cloud Transition's config-driven synthetic data platform. Its central architecture — a two-phase generation model that decouples row generation from foreign-key resolution — is a clean separation-of-concerns instance applied to relational data generation, paired with declarative simple/composite/polymorphic FKs, chunked (micro-batching) horizontal scaling, and percentile data-shape targets. The load-bearing lesson — distribution, not volume, is the bottleneck — is a data-skew argument for test-data design. The 2026-09-21 post is the canonical wiki instance.
  • Developer-productivity / CI-reliability at Atlassian-repo scale — Bitbucket Merge Queues defending against semantic merge conflicts at Jira-repo scale (800+ devs, 300+ merges/day, 70+ repos, 30,000+ PRs landed through merge queues). 2026-04-29 merge-queue architecture + outcomes post is the canonical wiki instance.
  • Jira Cloud multi-tenant configuration scaling — Jira Optimisation Tools paired with Site Optimiser and the Pre-computation Framework address configuration bloat at Atlassian-tenant-scale (100k+ users per tenant). The 2026-05-14 "Optimisation Tools for Jira" post is the canonical wiki instance, contributing the polymorphic usage tables decision (rejecting per-entity-type tables to avoid "millions of tables across production"), the Initialise → Scan-steps → Finalise batch-framework contract with idempotent + thread-safe + order-agnostic scan-step workers, the utilisation-prioritised refresh strategy (recompute only spaces near ≈70% of cap), the tiered Memcache + DB storage split, and the audit-log-as-rollback v1 contract.
  • Jira-as-substrate for AI-agent KTLO automation — the Atlassian Jira engineering team uses Jira's structured custom fields, workflow state machine, and workflow-transition automations as the substrate for delegating KTLO engineering chores to AI agents (likely Rovo Dev). The 2026-06-01 "How We Cut up to 80% of Engineering Chores Using AI Agents in Jira" post is the canonical wiki instance, contributing the work-item-as-agent-prompt framing, the status- transition-triggers-agent-with-custom-system-prompt mechanic, the daily-heuristic- cron-emits-agent-ready-work-items detection layer, the three-tier repo-specific → flag-for-creation → generic skill fallback chain, the per-test-category classify-then-dispatch mechanism, and the first-pass-investigator
  • human-merge-gate operational model. Reported outcomes: ~80% reduction in flaky-test eng hours (~1 engineering week saved per month); 500+ merged PRs in 70 days from stale-flag cleanup on the Jira repo. Composes load-bearingly with the Bitbucket-merge-queue substrate (systems/bitbucket-merge-queues) for the cleanup-throughput economics.

Key systems

  • systems/rovo-studio — Atlassian's automation-authoring surface; on the wiki it is the dispatcher in the feature-flag-cleanup Dispatcher → Coding Agent → Closer loop: a scheduled rule that queries Jira for eligible stale-flag tickets and triggers one coding-agent Bitbucket Pipelines run per ticket, fanning out in parallel.
  • systems/atlassian-unleash-pwa — production event PWA composition: static React/TypeScript client behind a global CDN; platform edge for routing/session/API-gateway duties; Jira Cloud as the workflow-shaped system of record; Jira Automation as scheduled control plane; and a thin no-database privileged service. Its client-consistency contract is explicit eventual readiness plus a durable local outbox and settled optimistic overlay; XP events are the source of truth and leaderboard rows are a rebuildable materialized read model. Reported: 500+ concurrent users, 24→16 startup calls, 8→0 duplicates, ~12 s→<1 s time to first useful content, and eight live production changes during the event.
  • systems/atlassian-localization-pipeline — localization throughput system for 20+ product languages. It pairs context-rich AI pretranslation with mandatory professional approval in Smartling, and moves internationalization correctness upstream through editor and code-review validation. AI-assisted remediation cleared 20,000+ legacy issues. The 2026-08-06 post reports FY25 H2 translation volume up 272% YoY, early TER from ~10% to <50% by language, and an estimated path to up to 50% cost savings through reviewer-edit feedback.
  • systems/smartling — translation management system hosting the pipeline's human approval gate; every AI draft, whether from an internal translator or vendor hub, requires professional translator sign-off before release.
  • systems/atlassian-long-horizon — Atlassian's reasoning-native agent orchestrator powering Rovo Chat. Single LLM, single context, up to 150-iteration reasoning loop. Replaced the hierarchical Hybrid Orchestrator. Flattened tool surface + progressive disclosure + context compaction + child instances + prompt layer ordering for prefix caching + adaptive reasoning.
  • systems/rovo-chat — Atlassian's user-facing AI assistant product; cross-product research and actions across Jira, Confluence, Bitbucket, JSM, Compass, and third-party connectors. Powered by Long Horizon.
  • systems/rovo-agent-harness — the distributed execution runtime beneath Long Horizon. Split-plane (control plane + on-demand compute sandbox); programmatic tool calling / code mode; dynamic model selection; callback bridge (no direct sandbox egress); CI-versioned warm-started tool package + active manifest; durable interactive async sub-agents; discovery-probe + checkpoint error recovery; human-approved self-evolution loop. Fireworks is the plausible (unnamed) sandbox substrate.
  • systems/atlassian-ml-studio — Atlassian's unified enterprise-scale ML development platform; three architectural pillars (composable versioned modules, workflow orchestrator with hot clusters + deterministic caching, embedded multi-layer compliance); serves ~120k monthly workflow runs across 100+ ML teams; backbone for Rovo Search/Chat, Teamwork Graph, Confluence AI.
  • systems/atlassian-jira-optimisation-tools — admin experiences + async backend workflows that compute per-space configuration-usage reports and apply bulk remediation actions (dissociate unused fields, split field configuration schemes, clean up unused work types). Powers Jira's response to the post-100k-user configuration-bloat problem.
  • systems/atlassian-jira-site-optimiser — admin dashboard surface that surfaces "configuration issues and improvement opportunities across spaces"; the human-facing entry point to the Optimisation Tools.
  • systems/atlassian-precomputation-framework — async batch-reporting platform on Atlassian's internal workflow orchestration engine; Initialise → Scan-steps → Finalise three-phase contract; idempotent + thread-safe
  • order-agnostic scan-step worker discipline; tiered Memcache + relational-DB storage with polymorphic usage tables. Canonical first-party wiki instance of patterns/async-projected-read-model.
  • systems/jira — host system / multi-tenant SaaS; recently extended for Jira Cloud limits + guardrails + optimisation-tools coverage; also the substrate for AI-agent KTLO automation via custom-fields + workflow-transition agent integration.
  • systems/atlassian-teamwork-graph — Atlassian's cross-product knowledge graph (Jira + Confluence + Bitbucket
  • …); named in the 2026-06-01 KTLO-AI-agents post as one of the three context substrates the agent reads (alongside the work item itself and the workflow-automation system prompt). Stub page; deeper architectural decomposition deferred to a future ingest.
  • systems/atlassian-fireworks — Firecracker-microVM orchestrator on Kubernetes; the "secure execution engine behind Atlassian's AI agent infrastructure." 100ms warm starts, live migration, eBPF network policy, shared volumes, snapshot filesystem restore, sidecar sandboxes; internal scheduler + autoscaler + Raft persistence + Envoy ingress. "Built in four weeks, entirely by LLMs."
  • systems/rovo-dev — Atlassian's AI development agent; deep Bitbucket + Bitbucket Pipelines integration; supports addressable skills (including multi-step orchestration meta-skills), prompt shortcuts like !review-pr for adversarial sub-agent review, and PR-bot participation on the PR itself.
  • systems/bitbucket-pipelines — Atlassian's CI/CD product; serves both as the automated quality gate Rovo Dev reads pipeline output from and as the execution substrate for the merge-queue pipeline (via the merge-queues: section in bitbucket-pipelines.yml).
  • systems/bitbucket-merge-queues — Atlassian's pre-merge validation queue for Bitbucket Cloud. 70+ repos in production across Jira, Rovo, Trello; 30,000+ PRs landed via merge queues since Beta; canonical first-party validate-against- future-state-of-main instance on the wiki.
  • systems/bitbucket — host of the repo, PRs, merge-queue admin surface, and Rovo Dev's operating surface.
  • systems/confluence-streaming-ssr — Confluence's React-18 streaming SSR pipeline; renderToPipeableStream + Suspense + NodeJS-transform state injection; ~40% FCP win.
  • systems/atlassian-streamhub — Atlassian's central event streaming platform (Kafka via MSK); 150B events/day, 3.2M events/sec peak. Multi-shard clusters with tiered storage, failover runbooks, companion regions, rate limiting, and client quotas. Migrated from Kinesis at 22B events/day.
  • systems/atlassian-events-rail — Atlassian's multi-tenant webhook evaluation/matching/delivery platform (10B events/month, ~6,000/sec sustained). Two-stage pipeline (CPU-bound matching → network-bound delivery) with layered fairness: negative-lookup cache, queue isolation per pipeline, per-tenant rate limiting, per-recipient concurrency limiting. Delivers to Forge apps, customer webhooks, and audit-log endpoints.
  • systems/data-brewery — Atlassian Cloud Transition's config-driven synthetic data generator platform. Produces SQL-dump datasets for Jira, Confluence, and ecosystem-app schemas from non-identifiable metadata + data shapes, for migration testing at enterprise scale. Central architecture is a two-phase generation model (generate all non-FK fields independently, then inject simple/composite/polymorphic FKs in dependency order using only existing parents) that decouples row generation from relationship resolution. Declarative relationships in config; four FK-injection components (Registry / ValueProvider / Injector / DependencyContext); chunked execution for bounded memory + fault isolation + horizontal scaling; percentile-shape targets applied at injection time. Open problems: schema drift and conditional field rules.

Key patterns / concepts

Localization pipeline and source-quality control (2026-08-06 axis)

  • translation-context-assembly — combine developer notes and UI semantics with translation memory, glossary terms, and locale style rules. This is bounded per-string context engineering, not unbounded prompt history.
  • localization-readiness — source-message property that eliminates English-specific string assembly, missing plural rules, and opaque descriptions before translation begins.
  • translation-edit-rate — human draft-edit distance as a per-language quality and cost signal, explicitly not a replacement for translator approval.
  • patterns/human-calibrated-llm-labeling — AI drafts every string while professional translators retain final sign-off in Smartling.
  • shift-left-i18n-validation — editor and pull-request checks keep source localization defects from entering translation; AI bulk fixes clear existing debt through developer-reviewed pull requests.
  • reviewer-edits-as-translation-learning-signal — approved edits improve later drafts while the human release gate remains intact.

Rovo Chat Long Horizon architecture (2026-06-18 axis)

  • progressive-tool-disclosure — exposing tool schemas on demand via meta-tools rather than loading all upfront; avoids per-product schema tax while keeping tools available.
  • concepts/context-compaction — trimming/summarising older tool outputs before each LLM call to keep long reasoning loops within token limits.
  • adaptive-reasoning-effort — calibrating reasoning depth to query complexity; simple lookups get fast answers, multi-step research gets deep planning.
  • concepts/context-engineering — ordering prompt layers from stable-to-volatile to maximise LLM provider prefix-cache reuse.
  • concepts/observability — hierarchical trace trees for debugging long-running LLM agent loops.
  • single-loop-agent-orchestration — one LLM, one context, one iterative loop replacing multi-agent hierarchy.
  • patterns/tool-surface-minimization — collapsing per-product sub-agents into typed namespaced actions called directly.
  • progressive-tool-disclosure-meta-tools — two meta-tools per namespace (get_schema + invoke_tool) for on-demand schema loading.
  • prompt-layer-ordering-for-cache-hits — assembling prompts from most-stable to most-volatile for cache maximisation.
  • context-compaction-service — dedicated service trimming/summarising older context before each LLM call.

Jira-as-substrate for AI-agent KTLO automation (2026-06-01 axis)

  • ktlo-engineering-chores — first-class wiki home for the "keeping the lights on" work-category framing. Atlassian names the canonical KTLO category list (flag cleanup, flaky tests, vulnerabilities, a11y fixes, long-tail bugs) and the pattern-recognition prerequisite argument for delegation: "That pattern recognition is what makes delegation to agents possible."
  • work-item-as-agent-prompt — Jira work item treated as a structured agent prompt; structured custom fields carry per-task brief; workflow state machine is the orchestration layer; transition automation carries the per-procedure system prompt; cross-product Teamwork Graph supplements with shared cross-task context.
  • agent-as-first-pass-investigator — operational model: agent does investigation + diagnosis + draft PR; human reviews + merges; "hours of manual investigation can now become minutes of review." Triage-vs-fix split with bounded false-positive comment-only exit.
  • jira-status-transition-triggers-agent-workflow — Jira-Cloud-native mechanic: status change triggers agent run with custom system prompt encoded into the transition automation. Single-source-of-truth lifecycle preserved on the work item.
  • agentic-pr-triage — daily cron scans codebase + cross-references compliance / experiment / release-track state, emits one Jira work item per stale flag with full pre-resolved structured fields (flag name + type + repo + paths + line numbers + desired final state). Detection layer for the agentic remediation pipeline.
  • agent-skill-with-fallback-chain — three-tier fallback for per-codebase variability: repo-specific cleanup skill → flag the repo as needing one + provide owner instructions → generic cleanup skill. The middle tier converts "no skill exists" into operational signal.
  • test-category-classifier-then-specialist-skill — for flaky-test triage / fix: classify the test as unit / integration / visual regression, dispatch to the matching specialist skill. CPU-throttled-loop reproduction discipline. Sibling skill-dispatch axis to the per-codebase fallback chain.
  • Reported outcomes: ~80% reduction in flaky-test eng hours (~1 engineering week saved per month); 500+ merged PRs in 70 days from stale-flag cleanup on the Jira repo alone.

Jira Cloud multi-tenant configuration scaling (2026-05-14 axis)

  • configuration-bloat — first-class wiki home for the multi-tenant SaaS failure mode where per-tenant configuration entities (fields / work types / schemes / workflows / screens) accumulate over time, slow read paths, and push tenants past hard caps.
  • polymorphic-usage-tables-multi-tenant — application-schema-layer decision to use a small fixed set of generic tables keyed by (tenant, entity_type, entity_id, …) rather than one dedicated table per entity type. Rejected alternative would create "millions of tables across production" in Jira's multi-tenant architecture. Application-layer cousin of catalog bloat.
  • utilization-prioritised-refresh-strategy — recompute pre-computed reports only for tenants near limits (≈70% utilisation); accept acknowledged staleness elsewhere; surface staleness state + on-demand refresh.
  • concepts/idempotent-operations — the load-bearing scan-step worker contract for orchestrator-led batch processing; the trio together enables the framework to "parallelise and retry without coordination."
  • patterns/async-projected-read-model — three-phase Initialise → Scan-steps → Finalise framework; canonical first-party wiki instance via Atlassian's Pre-computation Framework.
  • polymorphic-usage-tables-for-multi-tenant-scale — the multi-tenant DB-design pattern for the persistent storage tier.
  • prioritised-refresh-by-utilisation-threshold — the refresh-policy lever that controls compute budget; paired with last-refresh-timestamp + on-demand-refresh UX commitments.
  • tiered-state-management-memcache-plus-db — Memcached for short-lived intra-job state; relational DB for persistent reports.
  • audit-log-as-rollback-substrate — v1 recoverability contract for destructive bulk operations: targeted opinionated actions + comprehensive audit logging + manual rollback (with future assisted / guided-rollback layers).

Agentic development (2026-04-24 Fireworks axis)

  • concepts/micro-vm-isolation — the substrate shape Fireworks productises internally; structurally adjacent to systems/fly-kubernetes.
  • black-box-validation — "I test outputs, not read code." The process claim that makes "four weeks, entirely by LLMs" plausible.
  • ai-writes-own-tests — agent authors both production code and the e2e test suite; the test suite is the primary correctness proof.
  • agent-orchestration-skill — skills as multi-step procedural runbooks, not single-tool bindings. Fireworks team has a meta-workflow skill for end-to-end development plus a narrower dev-shard-lifecycle skill.
  • adversarial-review-persona — sub-agent with an adversarial prompt that red-teams PRs before human review.
  • concepts/agentic-development-loop — the closed LLM → execution-environment → feedback loop Rovo Dev implements.
  • concepts/blast-radius — five-lever safety net (CI, sharding, RBAC+JIT, progressive rollout, AI-written e2e tests).
  • ai-writes-own-e2e-tests — agent authors the e2e suite, deploys to a dev shard, loops on failures until green.
  • dev-shard-iteration-loop — every feature deploys to a per-developer isolated dev shard on a real K8s cluster.
  • patterns/specialized-agent-decomposition — !review-pr spins up an independent adversarial reviewer; canonical implementation.
  • agent-orchestration-meta-skill — codebase-specific multi-step procedural skill.
  • agentic-pr-triage — three-tier (adversarial sub-agent → CI gate → human architect) review stack.
  • three-workspace-parallel-agent-workflow — three checkouts, three branches, three agents, one human dispatcher.
  • agentic-pr-triage — agent reads Bitbucket Pipelines output and addresses failures before requesting review.
  • rbac-jit-as-agent-safety-net — access-control lever in the five-lever safety net.

Web-tier performance (2026-04-16 Confluence axis)

  • concepts/streaming-ssr — emit HTML progressively at Suspense boundaries
  • react-hydration — reuse server markup on the client without re-rendering; ordering constraints with streaming
  • concepts/backpressure — nginx proxy_buffering and compression middleware as streaming-hostile defaults
  • suspense-boundary — progressive-rendering unit
  • asset-preload-prediction — feedback-loop bundle preload to unblock hydration
  • patterns/staged-rollout — per-percentile A/B rollout with guardrail metrics

Developer-productivity / CI-reliability (2026-04-29 Bitbucket Merge Queues axis)

  • merge-queue — queue accepted PRs, validate against the target branch's future state, merge or eject. Atlassian is the canonical first-party wiki instance across 70+ repos.
  • semantic-merge-conflict — the load-bearing failure mode merge queues defend against. Jira-repo disclosure: 7–10% of CI failures pre-queue, near zero post-queue.
  • build-reliability — developer-satisfaction on build reliability lifted 70% → 82% on Jira after merge-queue rollout.
  • ci-reliability — target-branch CI-reliability as the load-bearing metric the merge-queue pattern optimises.
  • trunk-based-development — the branching model merge queues defend (Jira: single main, 800+ devs, 300+ merges/day).
  • developer-velocity — the composite outcome metric Atlassian uses to motivate the merge-queue investment.
  • validate-against-future-state-of-main — the canonical architectural pattern. Temporary merge-queue-* branch + dedicated merge-queue pipeline + merge-or-eject.
  • eject-failing-pr-keep-queue-running — the failure-recovery discipline that keeps the queue throughput stable. Failed build ejects the PR, not the queue.
  • parent-child-pipelines-for-ci-parallelism — Jira's merge-queue pipeline fans out to three parallel parent-child pipelines, one per product distribution.

Recent articles

  • 2026-09-24 — sources/2026-09-24-atlassian-how-we-automated-feature-flag-cleanup-with-agentic-pipelines (How we automated feature-flag cleanup with Agentic Pipelines — the feature-flag-cleanup instance of the Dispatcher → Coding Agent → Closer loop, sibling to the 2026-08-28 vulnerability-remediation post. A scheduled Rovo Studio rule queries Jira for eligible stale-flag tickets and fans out one coding-agent Bitbucket Pipelines run per ticket; the agent loads a feature-flag-cleanup skill, verifies the flag's state before editing (patterns/verify-before-changing-code), keeps the surviving branch, removes dead code/gate/imports at each call site, updates tests, runs checks, and opens a PR only if they pass — then comments/labels the ticket. Human reviews + merges. Load-bearing thesis: thin prompt, rich skill — the versioned skill, not the prompt, holds the domain logic. In production since April 2026, run monthly.)
  • 2026-09-21 — sources/2026-09-21-atlassian-how-atlassian-built-a-scalable-synthetic-data-engine (How Atlassian built a scalable synthetic data engine — Tier-3 source that passes scope because it documents a real platform architecture: a two-phase generation model, a config-driven engine design decision, chunked horizontal scaling, and a named FK-injection engine with concrete trade-offs. Atlassian's Cloud Transition org needed realistic synthetic Jira/Confluence datasets to validate on-prem→Cloud migration tooling before enterprise customers migrated; the prior API-driven generator took multiple weeks per test environment and hit API rate limits. New system: Data Brewery — a config-driven synthetic data platform whose central mechanic decouples row generation from foreign-key resolution (separation of concerns): Phase 1 generates all non-FK fields per table independently (parallelizable); Phase 2 injects simple/composite/polymorphic FKs in dependency order using only existing parent rows, applying the requested data shape as percentile targets (work-items-per-space min 100 · p50 1,000 · p90 8,000 · p99 50,000 · max 500,000). The FK injection engine has four components (ForeignKeyRegistry, ForeignKeyValueProvider, ForeignKeyInjector, DependencyContext — the last lets a child reference a same-run parent with no DB round-trip). Jobs are chunked for bounded memory + fault isolation + horizontal scaling (micro-batching discipline). Load-bearing lesson: volume is not the bottleneck — distribution (shape) is (data skew) — one space with thousands of GB of attachments stresses a migration pipeline differently than the same volume across many spaces. Enriched concepts/separation-of-concerns (two-phase-generation Seen-in), concepts/micro-batching (chunked bulk-generation Seen-in), concepts/partition-skew-data-skew (shape-as-test-data-design Seen-in), concepts/schema-evolution (new "schema drift in config-driven tooling" section), and systems/jira (new synthetic-data-target section). Applied the taxonomy gate: did not mint pages for "two-phase relational data generation", "config-over-code generation", or "percentile-shape test data" — single-source, article-specific ideas recorded as prose + tags and canonicalized into separation-of-concerns / micro-batching / partition-skew. Open problems: schema drift (config vs evolving DB schemas; startup validation now, schema-registry integration exploratory) and conditional field rules (values that depend on other generated values, not yet expressible in declarative config).) (Opening the Door to Agent Autonomy: The Architecture Behind Rovo's Agent Harness — Tier-3 source that passes scope because it documents the distributed execution architecture beneath Rovo Chat with real design trade-offs and operational numbers. The Rovo Agent Harness is the runtime layer under Long Horizon: a split-plane architecture separating a durable control plane (conversation + harness, one scaled service for millions of users, high agent density because agents live without a sandbox each) from a disposable, elastically-sized, on-demand compute sandbox — chosen over a monolithic in-sandbox design so a sandbox crash never kills the conversation (graceful degradation + split-control-plane-and-sandbox). In the sandbox it runs programmatic tool calling ("code mode") with progressive disclosure to expose thousands of Jira/Confluence/third-party actions — measured >50% latency reduction, 55% fewer tokens, +30% accuracy vs a meta-tool baseline on complex Jira queries. It does not always materialize a sandbox: dynamic model selection (dynamic-sandbox-vs-direct-tool-call-routing) sends lightweight queries to a stateless direct tool call and deep analysis to a stateful sandbox. Tools are shipped as a CI-versioned localized Python package applied on sub-100ms warm start with a runtime active manifest (ci-versioned-tool-package-warm-start); all tool calls (native or MCP) route through a callback bridge for permission validation + auth so the sandbox has no direct external network access (callback-bridge-for-sandbox-egress). Multi-agent coordination uses durable interactive sub-agents — each a hidden Rovo Chat conversation with a durable handle, async execution, optional shared sandbox, and a wake-the-parent notification loop — replacing the prior single-use synchronous sub-agents that forced re-hydration. Error recovery is first-class: discovery probes (discovery-probe-error-enrichment) return enriched structured feedback on generated-code errors for self-correction, and execution checkpointing enables precise rollback so a failure on step 10 of 11 avoids duplicate artifacts. A human-approved self-evolution loop feeds flagged production telemetry back into the eval suite. New system: systems/rovo-agent-harness. New concepts: split-plane-agent-architecture, concepts/model-first-routing, callback-bridge, discovery-probe-error-recovery, concepts/durable-execution, self-evolving-agent-harness. New patterns: split-control-plane-and-sandbox, dynamic-sandbox-vs-direct-tool-call-routing, ci-versioned-tool-package-warm-start, callback-bridge-for-sandbox-egress, interactive-async-sub-agent, discovery-probe-error-enrichment. Extended: systems/rovo-chat, systems/atlassian-long-horizon, systems/atlassian-fireworks (plausible unnamed sandbox substrate), code-generation-over-tool-calls (Atlassian instance + numbers), concepts/context-compaction. Caveat: the sandbox substrate is not named — the Fireworks link is an inference; code-mode gains are relative to an unspecified baseline; checkpoint granularity, manifest schema, sub-agent lifetime, and self-evolution workflow are qualitative.)

  • 2026-08-07 — sources/2026-08-07-atlassian-building-a-real-time-pwa-on-atlassians-own-stack (Building a real-time PWA on Atlassian's own stack — Tier-3 source that passes scope because it documents a production client/system architecture and its consistency/performance trade-offs. The Unleash PWA serves 500+ concurrent users using a static React/TypeScript PWA, Atlassian platform edge, Jira Cloud as workflow-shaped system of record, Jira Automation as control plane, and a thin stateless service for privileged writes, notification signing, and spike smoothing. Key resilience decision: readiness is eventual, not Boolean; the client persists intent in a durable local outbox, applies an optimistic overlay, marks successful writes settled, and retains the overlay until a live read confirms it (durable-client-outbox-with-settled-overlay). XP events are source-of-truth write records; leaderboard rows are a rebuildable materialized view maintained through near-real-time, periodic incremental, and full recomputation. Performance outcomes: page-load calls 24→16 and duplicates 8→0 through in-flight-promise sharing; add-to-schedule calls 11→5 with mutation-aware sparse revalidation; bundle ~7.7 MB→~1.53 MB after fixing minification; first useful content ~12 s→<1 s through local/bundled data plus SWR. Returning users skip idempotent provisioning after an access check, saving ~5 s. Eight features/bugs shipped live through Rovo Dev → Bitbucket → Pipelines → production. Caveats: platform-edge, client-store, retry, auth, quota, notification, and latency-SLO implementation details are undisclosed.)

  • 2026-08-06 — sources/2026-08-06-atlassian-scaling-localization-at-atlassian-keeping-translation-at-the-pace-of-ai-era-development (Scaling localization at Atlassian: keeping translation at the pace of AI-era development — Tier-3 source that passes scope because it documents a concrete production localization workflow and scaling trade-offs. AI-assisted development drove FY25 H2 translation volume 272% YoY across 20+ product languages. Atlassian's coupled response is: (1) context-rich AI pretranslation from an internal system or vendor hub using developer notes, similar approved translations, glossary terms, and locale style rules; (2) professional translator approval of every string in Smartling; (3) editor and code-review validation that shifts untranslated text, English-only composition, missing plural rules, and missing context back to the authoring point. Early Translation Edit Rate ranges from ~10% to <50% by language. Reviewer edits feed future improvement, with up to 50% savings estimated, but not yet reported as realized. AI-backed bulk remediation cleared 20,000+ legacy issues through developer-reviewed pull requests. New systems: systems/atlassian-localization-pipeline, systems/smartling. New concepts: translation-context-assembly, translation-edit-rate, localization-readiness. New patterns: patterns/human-calibrated-llm-labeling, shift-left-i18n-validation, reviewer-edits-as-translation-learning-signal. Caveats: internal model/vendor mechanics, trial methodology, TER by-language distribution, TMS API workflow, validation rule implementation, remediation precision, and realized savings are undisclosed.)

  • 2026-07-29 — sources/2026-07-29-atlassian-inside-the-events-rail (Inside the Events Rail: How Atlassian Delivers 10 Billion Webhooks a Month — Multi-tenant event evaluation/matching/delivery platform. 10B events/month, ~6,000/sec sustained, 18% MoM growth. Two-stage pipeline (CPU-bound matching → network-bound delivery), layered multi-tenant fairness (negative-lookup cache, queue isolation, per-tenant rate limiting, per-recipient concurrency limiting), per-surface configuration model, TCP-shaped backpressure, circuit breakers. New systems: systems/atlassian-events-rail, systems/forge. New concepts: concepts/tenant-isolation, selective-pub-sub. New patterns: two-stage-match-then-deliver-pipeline, queue-isolation-per-pipeline, per-tenant-rate-limiting, per-recipient-concurrency-limiting, per-surface-configuration-model, negative-lookup-cache.)

  • 2026-07-28 — sources/2026-07-28-atlassian-scaling-streamhub-transitioning-from-kinesis-to-kafka (Scaling StreamHub: Transitioning from Kinesis to Kafka for 145B Daily Events — Multi-year migration from Kinesis to MSK at 150B events/day. Covers tiered storage, six managed-service failure modes, hardening with sharded clusters, failover runbooks, companion regions, rate limiting, and client quotas. New systems: systems/atlassian-streamhub. New patterns: failover-cluster-escape-hatch, patterns/cell-based-architecture-for-blast-radius-reduction, ingress-rate-limiting-and-quarantine, patterns/async-replication-for-cross-region, conservative-broker-headroom.)
  • 2026-07-24 — sources/2026-07-24-atlassian-online-index-migration-and-shard-scaling-in-opensearch-with-aosc (Online index migration and shard-scaling in OpenSearch with AOSC — Atlassian Search Platform's open-source OpenSearch plugin for zero-downtime index reshaping. Applies CDC-style backfill+replay using Lucene changes snapshot and retention leases. Runs against 300+ clusters / 2500+ nodes / 1.7 PB. New systems: systems/opensearch-aosc. New concepts: retention-lease, lucene-changes-snapshot, index-alias. New patterns: patterns/shadow-migration, adaptive-rate-control, coordinator-worker-split.)
  • 2026-06-18 — sources/2026-06-18-atlassian-long-horizon-reasoning-engine (Long Horizon: How Atlassian Built a Reasoning Engine for Complex AI Tasks — first-party architectural overview of Rovo Chat's Long Horizon reasoning engine. Replaces the hierarchical Hybrid Orchestrator (per-product sub-agents) with a single-LLM, single-context, 150-iteration reasoning loop. Key architectural decisions: (1) flattened tool surface — every product's tools as typed namespaced actions called directly; (2) progressive disclosure via two meta-tools per namespace (get_schema + invoke); (3) context compaction service trimming older outputs with on-demand retrieval; (4) prompt layer ordering (stable-to-volatile) for prefix-cache maximisation; (5) child instances for parallel wide-task decomposition. Production results: +8.5% offline accuracy, +23% task completion (Confluence), −37% perceived latency. Canonical wiki instance of systems/atlassian-long-horizon, systems/rovo-chat, progressive-tool-disclosure, concepts/context-compaction, adaptive-reasoning-effort, concepts/context-engineering, concepts/observability, single-loop-agent-orchestration, patterns/tool-surface-minimization, progressive-tool-disclosure-meta-tools, prompt-layer-ordering-for-cache-hits, context-compaction-service. Caveats: no cost/token-usage numbers; compaction heuristics unspecified; adaptive reasoning mechanism undisclosed; no tool-selection accuracy empirical evidence for flattened surface.)
  • 2026-06-10 — sources/2026-06-10-atlassian-architecting-scalable-ml-platforms (Architecting Scalable ML Platforms: The Integrated Infrastructure and Acceleration Behind Rovo — first-party architectural overview of ML Studio, Atlassian's unified enterprise-scale ML development platform. Three architectural pillars: (1) composable reusable ML modules as versioned artifacts (2,000+ modules, 200k+ monthly iterations, <30s local builds); (2) workflow orchestrator with hot clusters, deterministic caching (~80% of workflows, 1,000+ hours/month saved), nested/joined workflows, CRON scheduling (~120k monthly runs, 100+ ML teams); (3) embedded multi-layer compliance (user identity → domain-level → column-level classification with automatic tag propagation, 900k+ governed datasets). Integration layer connects experiment tracking, central feature store, model registry, ML Lens monitoring, and serving platform. Serves models to 5M+ monthly active Rovo users. Canonical first-party wiki instance of systems/atlassian-ml-studio, ml-platform-architecture, workflow-orchestration, composable-ml-modules, concepts/fine-grained-authorization, automatic-task-caching, hot-cluster-reuse, concepts/centralized-ai-governance, module-as-versioned-artifact, hot-cluster-for-iterative-ml, deterministic-task-caching, column-level-classification-propagation, nested-composable-workflows, local-dev-loop-with-remote-parity. Caveats: no failure-handling / retry semantics; no cluster autoscaling or cost-management details; no GPU topology specifics; caching correctness guarantees unspecified.)
  • 2026-08-28 — sources/2026-08-28-atlassian-agentic-automation-in-practice-putting-standard-engineering-work-on-autopilot (Agentic automation in practice: putting standard engineering work on autopilot — the vulnerability-remediation companion to the 2026-06-01 KTLO post. Canonical wiki instance of the Dispatcher → Coding Agent → Closer three-part loop: each stage a versioned Bitbucket Agentic Pipeline step invoking Rovo Dev non-interactively, with scoped Atlassian MCP tool permissions per stage. The Dispatcher fetches/classifies/dedupes/ batches/dispatches; the Coding Agent invokes a fix-vulnerability decision-tree skill (direct dep bump / platform-baseline bump / transitive-parent bump / forced override / base-image tag / sidecar pin / "not our fix — close and flag") and opens a PR only on a green build (agentic-pr-triage); the idempotent Closer verifies the fix is deployed (not just merged) before transitioning the Jira item, leaving it open otherwise. Coins the wiki concept "the prompts are the system" — institutional knowledge as versioned, reviewable, in-repo artifacts — and agentic-vulnerability-remediation as the fix-loop workload class. Also canonicalises dedicated-service-account-for-agent-prs (bot identity for agent PRs). Reported: 120+ vulnerabilities resolved, 55+ automated PRs merged, 95% first-run merge rate (May–July 2026), "zero engineer effort beyond PR review." New concepts (2): agentic-vulnerability-remediation, prompts-as-the-system. New patterns (4): dispatcher-coding-agent-closer, decision-tree-skill, idempotent-closer-verifies-deployment, dedicated-service-account-for-agent-prs. Existing pages extended: systems/rovo-dev (vuln-remediation instance), systems/bitbucket-pipelines (Agentic Pipelines instance), systems/jira, systems/model-context-protocol, ktlo-engineering-chores (vulnerability category made concrete), agent-skill (decision-tree-skill shape), concepts/human-in-the-loop (PR-review gate), concepts/idempotent-operations (closer re-run safety). Caveats: Tier-3 product post; self-reported metrics on Atlassian's own repos; config snippets are excerpts; no breakdown of the 5% first-run failures.)
  • 2026-06-01 — sources/2026-06-01-atlassian-how-we-cut-up-to-80-of-engineering-chores-using-ai-agents-in (How We Cut up to 80% of Engineering "Chores" Using AI Agents in Jira — first-party Jira-engineering-team post on using AI agents to automate KTLO ("keeping the lights on") maintenance work. Canonicalises the KTLO engineering-chores category list (flag cleanup, flaky tests, vulnerabilities, a11y fixes, long-tail bugs) and the pattern-recognition-as-prerequisite argument for delegation: "Our team has spent years fixing these exact categories of issues. That pattern recognition is what makes delegation to agents possible." First wiki articulation of work-item-as-agent-prompt — Jira work item as structured agent prompt (custom fields = task brief; workflow state machine = orchestration; transition automation = custom system prompt; Teamwork Graph = cross-product context). Two production examples on the Jira repo: (1) flaky-test triage and fix with a test-category classifier dispatching to unit / integration / visual-regression specialist skills (test-category-classifier-then-specialist-skill) and CPU-throttled-loop reproduction; (2) stale feature flag cleanup via daily heuristic cron emitting fully-pre-resolved Jira work items (agentic-pr-triage) routed through Jira-status-transition-triggered agent runs (jira-status-transition-triggers-agent-workflow) using a three-tier repo-specific → flag-for-creation → generic skill fallback chain (agent-skill-with-fallback-chain). Operational model is agent does first pass, human does merge gate: agent does investigation + diagnosis + draft PR; "hours of manual investigation can now become minutes of review." Reported outcomes: ~80% reduction in flaky-test eng hours; ~1 engineering week saved per month; 500+ merged PRs in 70 days from stale-flag cleanup on Jira alone (~7 PRs/day on a single KTLO category). Composes load-bearingly with Bitbucket Merge Queues — the throughput claim only works because the merge queue prevents semantic-merge-conflicts that would otherwise compound from agent-driven PR volume. New systems (1): systems/atlassian-teamwork-graph (stub). New concepts (3): ktlo-engineering-chores, work-item-as-agent-prompt, agent-as-first-pass-investigator. New patterns (4): jira-status-transition-triggers-agent-workflow, agentic-pr-triage, agent-skill-with-fallback-chain, test-category-classifier-then-specialist-skill. Existing pages extended: flaky-test (KTLO-automation framing added), concepts/feature-flag (stale-flag-cleanup-as-KTLO axis added), agent-orchestration-skill (per-codebase-variability axis + skill-fallback chain), concepts/agentic-development-loop (KTLO axis as third source instance), agentic-pr-triage (Jira-substrate sibling), systems/jira (third altitude: workflow-substrate-for-agents), systems/rovo-dev (KTLO-automation use case). Source post does not explicitly name Rovo Dev as the agent; the integration surface is Atlassian's documented Add an agent to workflow transitions feature, which aligns with Rovo Dev. Caveats: no model / cost / token disclosure; no false-positive escape rate; no before/after PR-quality comparison; no agent-stuck recovery story; Teamwork-Graph access protocol / schema undisclosed.)
  • 2026-05-14 — sources/2026-05-14-atlassian-optimisation-tools-for-jira-reducing-configuration-bloat (Optimisation Tools for Jira: Reducing Configuration Bloat and Enhancing Performance — first-party architectural retrospective on Jira Cloud's response to multi-tenant configuration-bloat at 100k+-user scale. Names Jira's per-space limits (700 fields / 150 work types / 20,000 field options / 50 issue security levels / 50 grants), the two-phase reporting+remediation tool model, and three load-bearing architectural decisions: (a) the Pre-computation Framework with Initialise → Scan-steps → Finalise three-phase contract on top of Atlassian's internal workflow orchestrator, with scan-step workers required to be idempotent + thread-safe + order-agnostic so the orchestrator parallelises and retries without coordination; (b) polymorphic usage tables — application-schema-layer decision to use a small set of generic tables instead of one table per entity type, to avoid creating "millions of tables across production" in Jira's multi-tenant architecture, paired with Memcached for short-lived intra-job state under the tiered Memcache+DB storage split; (c) utilisation-prioritised refresh at ≈70% threshold rather than fixed-schedule recomputation, with last-refresh-timestamp + on-demand-refresh as the staleness UX commitments. Plus audit-log-as-rollback v1 contract: targeted opinionated bulk actions + audit log + manual rollback, with assisted-rollback already shipping and guided in-product rollback planned. Outcome numbers: tens of millions of unused fields and work types removed; some customers streamlining hundreds of thousands of fields and work types in a single day; vast majority of tenants kept within limits including largest enterprise customers. Canonical first-party wiki instance of patterns/async-projected-read-model, polymorphic-usage-tables-for-multi-tenant-scale, prioritised-refresh-by-utilisation-threshold, tiered-state-management-memcache-plus-db, and audit-log-as-rollback-substrate; first-class wiki home for configuration-bloat, polymorphic-usage-tables-multi-tenant, utilization-prioritised-refresh-strategy, and concepts/idempotent-operations.)
  • 2026-04-29 — sources/2026-04-29-atlassian-inside-atlassians-merge-queues (Inside Atlassian's Merge Queues: how we ship faster with fewer incidents — first-party architecture + production-outcomes post for Bitbucket Merge Queues. 70+ repos in production (Jira, Rovo, Trello, others); 30,000+ PRs landed through merge queues since Beta last quarter. Jira-repo scale disclosure: 800+ devs, 300+ merges/day; semantic-merge-conflict-caused CI failures 7–10% → near zero; incidents 3–5/week → rare edge cases; build time 40 min → 35 min; developer-satisfaction on build reliability 70% → 82%. Mechanism: temporary bitbucket-merge-queue-* branch + dedicated merge-queue pipeline defined via merge-queues: block in bitbucket-pipelines.yml + merge-commit strategy + build-concurrency = 14 + three parallel parent-child pipelines; failed builds eject the PR, not the queue. Canonical wiki instance of validate-against-future-state-of-main and eject-failing-pr-keep-queue-running; first-class name for semantic-merge-conflict and merge-queue on the wiki. Supersedes the 2026-04-16 open-beta launch post skip — this is the first-party architectural follow-up with real scale numbers.)
  • 2026-04-24 — sources/2026-04-24-atlassian-rovo-dev-driven-development (Rovo Dev Driven Development — how Atlassian's Fireworks team built a Firecracker-µVM-on-Kubernetes secure execution engine in four weeks using AI agents end-to-end; 100ms warm starts, live migration, eBPF policy, Raft-backed scheduler, Envoy ingress; three-workspace parallel-agent workflow, AI-written e2e tests, adversarial !review-pr sub-agent, orchestration meta-skill, dev-shard iteration on shared AWS scms K8s cluster; five-lever safety net: CI + sharding + RBAC/JIT + canary + AI-written e2e tests; "main deploy to dev without PRGB" as the load-bearing team-level process commitment.)
  • 2026-04-16 — sources/2026-04-16-atlassian-streaming-ssr-confluence (React 18 streaming SSR in Confluence: ~40% FCP win; NodeJS transform pipeline sequences state-before-markup per chunk; fixes intermediate-proxy buffering, buffer-mode regex cost, and React-18 context/hydration re-render bug)

Skipped (logged)

  • 2026-06-01 — The AI-native SDLC is paying off: 19% more PRs and 2-3 hours saved per developer per week (Tier-3 Rovo Dev productivity-measurement marketing post; quasi-experimental statistics on 3,400 repos + 6,200-developer survey; no distributed-systems internals, no scaling trade-offs, no infrastructure architecture, no production incidents — body is ~50% measurement framework + statistics tables, ~30% outcome tables, ~20% explicit Rovo Dev rollout-recommendation product pitch)
  • 2026-06-01 — From alert noise to action: how 24 Hour Fitness transformed IT Ops with JSM (Tier-3 customer-case-study product PR for Atlassian Service Collection)
  • 2026-05-14 — Inside Reddit's IT playbook: Building for scale and AI-readiness (Atlassian customer-success / webinar-recap; corporate IT modernization narrative + product pitch, zero architectural content)
  • 2026-05-13 — 3 ways AI alert grouping is transforming on-call engineering (Jira Service Management AIOps product PR; listicle format)
  • 2026-04-16 — Enhancing Developer Workflow with Rovo Dev and Notifications (product PR)
  • 2026-04-16 — Merge Queues for Bitbucket Cloud, now in open beta (product PR — superseded by the 2026-04-29 first-party architectural follow-up which is ingested as sources/2026-04-29-atlassian-inside-atlassians-merge-queues)
  • 2026-04-16 — Rovo Dev in Frontend Platform Engineering (AI-agent codemod tooling)
Last updated · 766 distilled / 2,225 read