Opening the Door to Agent Autonomy: The Architecture Behind Rovo's Agent Harness¶
Summary¶
Atlassian's Rovo Engineering describes the agent harness behind Rovo Chat's move from a synchronous data-retrieval chatbot (50+ third-party sources) into an "always-on sidekick" that plans and runs large, cross-cutting workflows. The rethink of the execution model rests on five architectural pillars: a split-plane architecture separating the control plane (conversation + agent harness) from an isolated compute sandbox (data plane); programmatic tool calling ("code mode") in that sandbox with progressive disclosure to expose thousands of Jira / Confluence / third-party API actions; dynamic model selection that decides per-interaction whether to run a cheap stateless direct tool call or spin up a stateful sandbox; a callback bridge that routes all sandbox-originated tool calls back through the control plane for auth / permission validation (so the sandbox needs no direct external network access); asynchronous multi-agent coordination via durable, pollable interactive sub-agents; robust error recovery (discovery probes + checkpointing); and systematic self-evolution through an integrated, human-approved evaluation pipeline. This is the harness/runtime layer that sits beneath the Long Horizon reasoning engine.
Key takeaways¶
- Split-plane, not monolith. Atlassian rejected placing the agent harness inside the sandbox (a monolithic design) and instead separated the control plane (conversation state + agent harness, the interface the user connects to) from the sandbox (isolated compute / data plane), so each scales, degrades, and recovers independently. "When the sandbox fails, the conversation should live on." If the sandbox crashes, the user connects to the control plane and the agent transparently retries and spins up a new sandbox container. (Source: sources/2026-08-27-atlassian-opening-the-door-to-agent-autonomy-the-architecture-behind-rovos-agent-harness)
- Agent density is the scaling win. Because the control plane is a single scaled service handling events for millions of users, each user's agent can respond at any moment as a background/always-on operator without first creating a sandbox instance — sandboxes are provisioned on demand, increasing the density of concurrently-live agents. The same split lets local chat clients share the architecture "without having to rely on a heavy local binary to run the agent" — agent improvements ship server-side, no manual upgrade.
- Independent + elastic sandbox sizing. The sandbox scales independently of the agent: heavy tasks get a larger sandbox instance; simple operations start on a smaller instance that is upgraded only if required compute increases (both vertical and horizontal scale-out).
- Programmatic tool calling ("code mode") beats sequential tool calls. Deep-analysis questions (parsing massive volumes of Jira work items, Atlas goals, external data) are answered by offloading data iteration and pagination to code execution in the sandbox rather than having the LLM sequentially invoke tools and handle pagination. Atlassian explicitly aligns this with Anthropic ("computer use / code mode") and Cloudflare ("programmatic tool calling"). Combined with progressive disclosure, it "can safely expose thousands of API actions" across first- and third-party products.
- Measured code-mode wins: internal evals found code mode reduced latency of complex Jira-related queries by over 50%, cut token consumption by 55%, and improved accuracy by 30% — measured against a baseline harness that used standard meta-tools instead of code execution across a representative set of complex queries.
- Dynamic model selection — don't always pay for a sandbox. For lightweight informational queries the control plane calls tools directly (a "quick, stateless direct tool call"); deep data analysis gets the "robust, stateful sandbox environment." The agent dynamically decides the most efficient execution path per interaction.
- Dynamic + static tool loading with sub-100ms warm starts. The CI/CD pipeline dynamically versions and packages tools into a localized Python package that updates the sandbox on a warm start, "maintaining sub-100ms startup times." At runtime the agent passes an active manifest to the sandbox indicating which APIs are enabled / disabled based on the user's context, knowledge-source selection, and skill enablement. The hybrid design gives the sandbox programmatic access to both statically-defined internal tools and third-party MCP servers.
- Callback bridge — sandbox has no direct external network access. Tools may be native (part of the agent) or hosted on a remote MCP server; in either case all tool calls route through Atlassian's internal server, where user permissions are validated and requests are authenticated before any external API is invoked. The bridge acts as a "secure translator": sandbox code executes a tool call → bridge routes the request back through the control plane to fetch data from external MCP servers → "preventing the isolated sandbox environment from needing direct external network access."
- Interactive sub-agents — async, durable, pollable. The prior design spawned single-use synchronous sub-agents with no way to detach or persist state, forcing the parent to re-initialize and re-hydrate a new instance for any follow-up. The new interactive sub-agents let a parent delegate independent work to background processes without blocking its own execution loop: a sub-agent can outlive the parent's response turn, maintain its own independent conversation history, be polled for status or receive follow-up instructions, and wake the parent on completion. Architecturally a sub-agent is not a lightweight ephemeral function call — it is instantiated as a hidden Rovo Chat conversation backed by a durable handle, asynchronous task execution, optional shared sandbox access, and a dedicated notification loop back to the parent.
- Error recovery is a first-class feature. "Intelligent and successful error recovery is a large portion of what makes an agent effective." Two mechanisms: Discovery Probes — the harness intercepts runtime errors (wrong function signature, unexpected param) and returns rich, structured execution feedback (valid function signature, available parameters, context) so the agent self-corrects instead of burning tokens in a blind retry loop; and Checkpointing — snapshotting execution state at key milestones enables precise rollbacks so a failure on step 10 of 11 doesn't re-run the whole script and create duplicate artifacts / "Entity Already Exists" errors.
- Systematic self-evolution via integrated evals. An agent harness goes stale as data shapes and user behavior shift, so the harness is deeply integrated into an automated evaluation pipeline. On a failed query or inefficient execution path, telemetry is flagged, analyzed, and automatically fed back into the eval suite — with engineers reviewing and approving optimizations before they're applied (human-guided, not yet fully autonomous). Atlassian frames this flywheel as building for a future where engineers focus on "defining goals, setting guardrails, guiding the agent evolution" over repetitive deterministic logic.
Operational numbers¶
| Metric | Value |
|---|---|
| Third-party data sources supported (Rovo Chat) | 50+ |
| Users served by the (single scaled) control plane | Millions |
| Code-mode complex-Jira-query latency reduction | >50% |
| Code-mode token-consumption reduction | 55% |
| Code-mode accuracy improvement | 30% |
| Sandbox warm-start time (dynamic tool package) | Sub-100 ms |
Systems / concepts / patterns extracted¶
Systems. Rovo Agent Harness (new, canonical); Rovo Chat (the product this harness powers); Long Horizon (the reasoning engine that runs on top of the harness); Fireworks (the Firecracker-microVM-on-Kubernetes secure-execution substrate that plausibly backs the sandbox — the post does not name it, so this is an inferred substrate link); MCP (third-party tool servers reached through the callback bridge); Code Mode (Cloudflare's named instance of the same programmatic-tool-calling idea).
Concepts. Split-plane agent architecture (new), a specialization of control-plane / data-plane separation; dynamic model selection for agents (new); callback bridge (new); discovery-probe error recovery (new); agent execution checkpointing (new); self-evolving agent harness (new); plus existing progressive tool disclosure, graceful agent degradation, context window as token budget.
Patterns. split control plane and sandbox (new); dynamic sandbox vs direct tool-call routing (new); CI-versioned tool package + warm-start (new); callback bridge for sandbox egress (new); interactive async sub-agent (new); plus existing code generation over tool calls and checkpoint before risky step.
Caveats¶
- The post does not name the sandbox substrate; the Fireworks link is an inference from Atlassian's own prior disclosure that Fireworks is "the secure execution engine behind Atlassian's AI agent infrastructure," not a stated fact in this article.
- No absolute latency/throughput numbers, sandbox instance sizes, or concurrency figures are given; the code-mode gains are relative to an unspecified meta-tool baseline on a "representative set" of complex queries.
- Checkpointing granularity ("key milestones"), discovery-probe coverage, the manifest schema, sub-agent scheduling / lifetime limits, shared-sandbox isolation semantics, and the self-evolution approval workflow are described qualitatively only.
- Self-evolution "still requires human input" — the fully-autonomous flywheel is aspirational.
Source¶
- Original: https://www.atlassian.com/blog/rovo/agent-autonomy
- Raw markdown:
raw/atlassian/2026-08-27-opening-the-door-to-agent-autonomy-the-architecture-behind-r-ba8b0a18.md
Related¶
- systems/rovo-agent-harness
- systems/rovo-chat
- systems/atlassian-long-horizon
- split-plane-agent-architecture
- concepts/model-first-routing
- callback-bridge
- discovery-probe-error-recovery
- concepts/durable-execution
- self-evolving-agent-harness
- split-control-plane-and-sandbox
- dynamic-sandbox-vs-direct-tool-call-routing
- ci-versioned-tool-package-warm-start
- callback-bridge-for-sandbox-egress
- interactive-async-sub-agent
- companies/atlassian