Skip to content

ATLASSIAN 2026-09-24 Tier 3

Read original ↗

How we automated feature-flag cleanup with Agentic Pipelines

Summary

A first-party Atlassian Engineering (Bitbucket blog) post describing a production agentic pipeline that clears a team's monthly backlog of stale feature flags. It is the feature-flag-cleanup instance of the same Dispatcher → Coding Agent → Closer architecture Atlassian documented for vulnerability remediation a month earlier (2026-08-28). A scheduled Rovo Studio automation rule queries Jira for stale flag-cleanup tickets that are eligible for automated removal and, for each one, triggers a coding-agent Bitbucket Pipelines run with the ticket key passed as context — fanning out across every eligible ticket in parallel. The agent reads the ticket, loads a feature-flag-cleanup skill, verifies the flag's current state in the codebase before touching anything, inlines the surviving behaviour, deletes the dead branch + gate definition, updates affected tests, runs the repo's required checks, and — only if checks pass — opens a pull request, comments the PR link on the ticket, and labels it as automated. An engineer still reviews and merges. In production since April 2026 on one Atlassian team, run monthly. The load-bearing thesis: keep the prompt thin and the skill rich — almost none of the domain logic lives in the prompt; it lives in a versioned, reviewable skill file. (Source: How we automated FF cleanup with Agentic Pipelines)

Why this problem

  • The hard part of a feature flag is removing it, not adding it. By the time a flag is safe to delete, "the rollout is over, the original context has faded, and there is always a more urgent piece of work waiting." Cleanup tickets perpetually lose to the backlog.
  • A growing backlog is a tax on everyone who touches the surrounding code. "Engineers have to reason around those flags whenever they touch the surrounding code." A few stale flags are fine; a growing pile creates confusion.
  • The steps are known but not scriptable. Find every usage, inline the final behaviour, remove the dead branch, clean up definitions, run tests, open a PR — "most teams already have this written down somewhere." But it "requires codebase-specific knowledge that isn't easy to automate with a script," which is exactly what makes it a fit for an agent loaded with a codebase-specific skill rather than a static script.

The architecture: dispatcher + coding agent

The automation has two parts: a dispatcher that finds the work, and a coding agent that does it — the same shape as the Dispatcher → Coding Agent → Closer loop, with the closer responsibilities (comment PR link, label ticket) folded into the tail of the coding-agent run here.

The flow, verbatim from the post:

Rovo Studio automation rule (scheduled)
  → queries Jira for stale flag-cleanup tickets ready for code removal
  → for each ticket: triggers a coding-agent pipeline run (ticket key passed as context)
      → agent reads ticket metadata (flag name, final value)
      → loads the flag-cleanup skill
      → searches the codebase — verifies current state before touching anything
      → inlines the final behavior, removes the dead code path
      → runs linters, type-checks, and tests
      → opens a pull request
      → comments on the ticket with the PR link and labels it as automated
  → engineer reviews and merges the PR
  → flag is archived in the team's flag service post-deploy → ticket closes automatically

The dispatcher: the work finds the agent

A scheduled Rovo Studio automation rule queries Jira for open flag-cleanup tickets that are stale and eligible for automated removal. For each one, it triggers the coding agent as a Bitbucket Pipelines run, passing the ticket key as context.

"Nobody has to remember to invoke anything: the work finds the agent. The dispatcher fans out across however many eligible tickets exist and runs agents in parallel."

This is the parallel fan-out over a queue of eligible work items property — the dispatcher removes the repeated manual startup work (finding stale flags, kicking off runs) and scales to N tickets by running N agents concurrently.

The coding agent: thin prompt, rich skill

The prompt is kept deliberately thin: the pipeline reads the ticket, loads a flag-cleanup skill, executes it, and reports a one-line result. "Almost none of the domain logic lives in the prompt itself. Everything lives in the skill." This is the same "prompts are the system / thin-prompt-rich-skill" thesis as the vulnerability-remediation post — an application of separation of concerns: don't scatter flag-cleanup knowledge across pipeline config and prompt text.

A simplified bitbucket-pipelines.yml from the post shows the agent definition + per-step OAuth scopes:

# bitbucket-pipelines.yml (excerpt)
definitions:
  agents:
    flag-cleanup:
      prompt: ".claude/flag-cleanup-agent.md"
      config:
        path: .claude/pipeline-config.yml
pipelines:
  custom:
    flag-cleanup:
      - step:
          name: Remove flag and open pull request
          auth:
            system:
              scopes:
                - read:repository:bitbucket
                - write:repository:bitbucket
                - write:pullrequest:bitbucket
          script:
            - agent: flag-cleanup

Key takeaways

  1. Verify before changing code — verification is step one, not a safeguard. "A cleanup ticket is useful context, not proof that the codebase is unchanged." Repos move fast; devs frequently touch adjacent code, so the repo state can diverge from the initial rollout plan — sometimes the flag was already removed, sometimes its final value changed. The agent searches for the flag (name and value) before editing; if nothing is found it stops, because the flag may already be gone. (See patterns/verify-before-changing-code.)
  2. Keep the prompt thin and the skill rich. The agent prompt should be a small harness that loads a skill and reports back — not a second place where domain knowledge lives. "That way the same skill works whether it's triggered by the pipeline or run interactively by a developer." The skill is "the part we keep rewriting."
  3. Let checks gate the pull request — do not skip past failures. "Checks have to pass before a pull request opens." If linting, type checks, or tests fail, the workflow investigates and repairs within its defined scope; if it cannot safely produce a passing change, it stops and reports rather than opening a broken PR. (See patterns/tests-as-executable-specifications.)
  4. Scope permissions to the task. "The agent only needs to read and write the specific ticket system and open pull requests. Task-specific auth tokens limit the damage if a run goes wrong." Repositories are configured with branch restrictions and policies to prevent agents from pushing directly to main; token scopes grant only the access each task requires. (See concepts/least-privileged-access.)
  5. Treat the skill as a living document. "Codebase conventions drift. A skill that isn't revisited after real runs will eventually miss an edge case that a developer would have caught by eye." The skill is versioned and reviewable — "when the patterns in the codebase change, the skill gets updated, and the next run uses the new version."
  6. Humans stay in the loop at review, not at startup. "The goal was not to remove engineers from the process. It was to remove the repetitive start-up work: finding eligible tickets, reconstructing a flag's history, locating every usage, and preparing the first safe change. Engineers still make the final call in pull-request review." (See concepts/human-in-the-loop.)
  7. The pattern generalizes to bounded, well-triggered maintenance work. "This pattern works best for maintenance work with reliable ticket context, a bounded code change, and clear validation." Feature flags met those conditions; other teams may find the same pattern useful for dependency updates, documentation fixes, or remediation work. Explicitly cross-linked to Atlassian's prior vulnerability-fixes and documentation agentic-pipeline posts.

The skill (shortened, from the post)

# Feature-flag cleanup
Read the ticket for flag name and final value. Search the repo for both
before editing. If nothing is found, stop — the flag may already be gone.

Keep the branch that matches the final value. Delete the other path,
the flag helper, and unused imports. Do this at each call site; the
pattern is often not uniform.

{isEnabled() ? <NewUI /> : <OldUI />}  →  <NewUI />
if (checkGate("flag")) next() else old()  →  next()

Update tests that mocked the discarded value. Run the repo's lint,
type-check, and tests. No passing checks, no pull request.

If the final value is false, or the surviving branch is unclear, stop
and ask. Archive the flag in the flag service after merge, not here.

Two properties worth noting:

  • Per-call-site handling. "The pattern is often not uniform" — the skill enumerates concrete transformation shapes (ternary → surviving component; gate conditional → surviving branch) and applies them at each site rather than a single blanket edit.
  • Ambiguity → stop and ask. "If the final value is false, or the surviving branch is unclear, stop and ask." Removing the enabled branch (final value false) or an unclear survivor is exactly where a wrong automated edit is most dangerous, so those cases escalate to a human rather than guess.

Operational numbers / facts

  • In production since April 2026 on one Atlassian team; run to clear the monthly stale-flag backlog. (Vendor-reported; no PR-count or time-saved figure disclosed in this post — unlike the sibling 2026-06-01 Jira-team post's "500+ merged PRs in 70 days.")
  • Ticket source: Atlassian's internal feature-flag gating system auto-creates one Jira ticket per flag; each stale-flag ticket carries flag name, final value, flag type, and the gate's current state.
  • Post-merge: the flag is archived in the team's flag service after deploy, which auto-closes the ticket.
  • Guardrails: branch restrictions prevent direct pushes to main; per-step OAuth scopes (read:repository, write:repository, write:pullrequest); checks must pass before a PR opens.
  • Agentic Pipelines is in beta for Bitbucket Cloud at time of writing.

Caveats / what's not disclosed

  • No throughput or savings numbers in this post (flags-cleaned/month, hours saved, first-run merge rate). The architecture and guardrails are the contribution; the quantified ROI lives in the sibling posts.
  • "Rovo Studio" vs "Rovo Dev" framing. The dispatcher is named as a Rovo Studio automation rule; the coding agent is invoked via a Bitbucket Agentic Pipeline whose .claude/flag-cleanup-agent.md prompt
  • pipeline-config.yml config match the Rovo Dev non-interactive-pipeline shape documented on 2026-08-28. The post does not name the underlying coding-agent model or MCP tool allowlist here.
  • Single team, one workload. The post is explicit that "there are variations to this in other teams based on their code base structure and rollout rules" — this is one team's instance, not a fleet-wide standard.

Source

Last updated · 766 distilled / 2,225 read