SYSTEM Cited by 1 source
Rovo Agent Harness¶
The Rovo Agent Harness is the distributed execution runtime beneath Rovo Chat that turns it from a synchronous data-retrieval chatbot into an "always-on sidekick" capable of planning and running large, cross-cutting workflows in the background. Where Long Horizon is the reasoning engine (the single-LLM iterative loop that decides what to do), the harness is the runtime that executes those decisions: it owns the split-plane topology, the compute sandbox, tool loading, the callback bridge, multi-agent coordination, error recovery, and the self-evolution loop.
Architecture: split plane¶
The harness deliberately separates the control plane from the data plane rather than embedding the agent inside the sandbox (a rejected monolithic design):
- Control plane — holds conversation state and the agent harness; it is the interface the user connects to. A single scaled service handling events for millions of users.
- Sandbox (data plane) — isolated compute where generated code runs. Provisioned on demand, sized elastically, and disposable.
This is a specialization of control-plane / data-plane separation for agents; see split-plane agent architecture and split control plane and sandbox. The design goals it buys:
- Graceful degradation. "When the sandbox fails, the conversation should live on." Sandbox crashes are recoverable with no user-perceived impact — the user is connected to the control plane, and the agent transparently retries and spins up a new sandbox container. A form of graceful agent degradation at the infrastructure layer.
- Efficient + independent scaling. The sandbox scales independently of the agent — heavy tasks get a larger instance; simple ops start small and upgrade only if compute demand grows (vertical + horizontal).
- Agent density. Because the control plane is always up, each user's agent can respond at any moment as a background operator without first creating a sandbox, maximizing the density of concurrently-live agents.
- Multi-platform support. Local chat clients share the same architecture "without having to rely on a heavy local binary to run the agent" — improvements ship server-side.
Programmatic tool calling in the sandbox ("code mode")¶
Deep-analysis questions (parsing large volumes of Jira work items, Atlas goals, external data) are answered by offloading data iteration and pagination to code execution in the sandbox rather than making the LLM sequentially invoke tools and hand-paginate — the code-generation-over-tool-calls pattern, explicitly aligned by Atlassian with Anthropic's "computer use / code mode" and Cloudflare's "programmatic tool calling" (Code Mode). Combined with progressive disclosure, the harness "can safely expose thousands of API actions" across first- and third-party products. The sandbox also serves as a place to store temporary information, generate visualizations, and create artifacts.
Measured against a meta-tool baseline on complex Jira queries, code mode cut latency >50%, token consumption 55%, and improved accuracy 30%.
Dynamic model selection¶
The harness does not spin up a sandbox for every interaction. Via dynamic model selection / dynamic sandbox vs direct tool-call routing, the agent picks the cheapest sufficient execution path per interaction: a stateless direct tool call from the control plane for lightweight informational queries, or a stateful sandbox environment for deep data analysis.
Dynamic + static tool loading¶
The CI/CD pipeline dynamically versions and packages tools into a localized Python package that updates the sandbox on a warm start, maintaining sub-100ms startup times (CI-versioned tool package + warm-start). At runtime the agent passes an active manifest to the sandbox indicating which APIs are enabled / disabled based on the user's context, knowledge-source selection, and skill enablement. The hybrid design gives the sandbox programmatic access to both statically-defined internal tools and third-party MCP servers.
Callback bridge¶
All tool calls — native or remote-MCP-hosted — route through Atlassian's internal server, where user permissions are validated and requests authenticated before any external API is invoked. The bridge is a "secure translator": sandbox code executes a tool call → the bridge routes it back through the control plane to fetch data → the isolated sandbox never needs direct external network access. See callback bridge / callback bridge for sandbox egress.
Agent coordination: interactive sub-agents¶
The harness supports a true asynchronous multi-agent system via interactive sub-agents ( interactive async sub-agent). A parent delegates independent work to a background process without blocking its own loop; the sub-agent can outlive the parent's turn, keep its own conversation history, be polled for status or given follow-up instructions, and wake the parent on completion. Architecturally a sub-agent is not an ephemeral function call — it is a hidden Rovo Chat conversation backed by a durable handle, async task execution, optional shared sandbox access, and a notification loop back to the parent. This contrasts with the prior design's single-use synchronous sub-agents that had no way to detach or persist state.
Error recovery¶
- Discovery probes. The harness intercepts runtime errors from generated code (wrong function signature, unexpected param) and returns rich, structured execution feedback — valid signature, available params, required context — so the agent self-corrects instead of looping blindly. See discovery-probe error recovery / discovery probe error enrichment.
- Checkpointing. Execution state is snapshotted at key milestones so a failure on step 10 of 11 rolls back precisely rather than re-running the whole script and creating duplicate artifacts / "Entity Already Exists" errors — see agent execution checkpointing, a logical sibling of checkpoint before risky step.
Self-evolution¶
The harness is deeply integrated into an automated evaluation pipeline: on a failed query or inefficient path, telemetry is flagged, analyzed, and fed back into the eval suite, with engineers reviewing and approving optimizations before they're applied. See self-evolving agent harness and its relation to the continuous evaluation feedback loop.
Relationship to Long Horizon¶
Long Horizon and the harness are complementary layers of the same Rovo Chat stack:
- Long Horizon = reasoning engine (one LLM, one context, up-to-150-step loop, flattened tool surface, prompt-layer ordering, adaptive reasoning).
- Rovo Agent Harness = distributed runtime (split plane, sandbox, code mode, dynamic model selection, callback bridge, interactive sub-agents, error recovery, self-evolution).
Concepts they share — progressive tool disclosure, context compaction, and sub-agent decomposition — are described from the reasoning side in Long Horizon and from the runtime side here.
Substrate note¶
The 2026-08-27 post does not name the sandbox substrate. Atlassian's own Fireworks is described elsewhere as "the secure execution engine behind Atlassian's AI agent infrastructure" (Firecracker microVMs on Kubernetes, 100ms warm starts, snapshot restore, sidecar sandboxes), which lines up with the harness's on-demand, elastically sized, warm-starting sandbox — so Fireworks is the plausible substrate, but this is an inference, not a stated fact.
Seen in¶
- sources/2026-08-27-atlassian-opening-the-door-to-agent-autonomy-the-architecture-behind-rovos-agent-harness — canonical architectural description of the harness.
Related¶
- systems/rovo-chat — the product the harness powers
- systems/atlassian-long-horizon — the reasoning engine on top of the harness
- systems/atlassian-fireworks — plausible microVM sandbox substrate
- systems/model-context-protocol — third-party tool servers reached via the callback bridge
- systems/code-mode — Cloudflare's named instance of programmatic tool calling
- split-plane-agent-architecture
- concepts/model-first-routing
- callback-bridge
- discovery-probe-error-recovery
- concepts/durable-execution
- self-evolving-agent-harness
- split-control-plane-and-sandbox
- dynamic-sandbox-vs-direct-tool-call-routing
- ci-versioned-tool-package-warm-start
- callback-bridge-for-sandbox-egress
- interactive-async-sub-agent
- code-generation-over-tool-calls