PATTERN Cited by 19 sources
Specialized agent decomposition¶
Build per-domain agents (storage, databases, client-side traffic, network, …) that each carry a small, well-scoped toolset, and let them collaborate on an end-to-end analysis — rather than building one mega-agent that carries every tool and context for every domain.
Intent¶
A single general-purpose agent suffers two failure modes as it grows:
- Tool-selection noise. Large tool inventories make the LLM more likely to pick wrong / less-optimal tools.
- Context crowding. Packing domain-specific system prompts into one context dilutes each domain's instructions and hits context-window limits.
Decomposition puts each domain's tools and prompts in a dedicated agent whose reasoning space is small, then composes their outputs for cross-domain investigations.
When to reach for it¶
- You already have a patterns/specialized-agent-decomposition: adding an agent ≈ adding a configuration, not a new codebase.
- Debugging / investigation spans multiple subsystems (e.g., DB + client traffic + storage).
- You observe tool-selection errors correlated with tool-inventory growth.
Mechanism¶
- Carve along coherent domains. Each agent owns a specific scope: one system-and-database agent, one client-traffic agent, etc. Tools within an agent are cohesive.
- Shared infrastructure. Framework (LLM client, conversation state, tool-call parser, snapshot/replay harness) lives once; each agent instantiates it.
- Collaboration protocol. Either an orchestrator agent routes questions to specialists and merges outputs, or specialists hand off to each other via well-defined events. Databricks' post describes collaboration but doesn't spec the protocol.
- Per-agent evaluation. Each agent has its own snapshot-replay corpus (see patterns/snapshot-replay-agent-evaluation); specialization makes eval more tractable, not less.
Why it helps¶
- Deep expertise per agent. Smaller tool inventory + focused prompt + focused eval corpus = better domain accuracy.
- Parallel team development. Different teams can own different specialist agents.
- Incremental rollout. New domains get their own agent without destabilizing existing ones.
- Extensibility beyond original scope. Once a few agents exist, adding one for a new system (say, caching, or Kubernetes) is a well-defined template.
Tradeoffs¶
- Orchestration overhead. Cross-domain questions now require coordination — "this is a DB issue triggered by a client-side surge" requires both agents. Poorly designed coordination layers regress latency and UX.
- Consistency. Multiple specialists can return overlapping or contradictory diagnoses. Need a reconciliation step or primary-agent mechanic.
- Boundary drift. A signal that looks like a DB issue may actually live in client traffic; agents must know when to hand off.
Seen in¶
-
sources/2026-10-01-aws-accelerating-airline-retailing-innovation-datalex-modernization — Three specialized agents behind one orchestrator for a booking assistant. Datalex's agentic-AI proof of concept uses a Bedrock AgentCore orchestrator coordinating authentication, data-retrieval, and reporting agents, each calling existing REST APIs through the AgentCore Gateway — the agents layer onto the modernized system without rewriting it.
-
sources/2026-09-30-aws-how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock — Controller-agent split with a config-driven agent registry. MHK's SmartProminence AI Orchestrator enforces a strict separation between workflow orchestration and LLM processing: a stateless Workflow Engine Controller resolves a DAG (Kahn's algorithm) and decides what runs and in what order; small, single-purpose LLM processing agents execute individual steps and operate only on their specific, actionable input. Each agent carries a focused prompt + input/output schema + model selection — the "small, well-scoped reasoning space per agent" thesis of this pattern, applied as an orchestrator/worker (controller-agent) split rather than domain-carved specialists. Adding an agent is a configuration (dynamic agent registry: define prompts/schemas, register via API, Terraform provisions the SQS queue + IAM role + ECS task) — the "adding an agent ≈ adding a configuration" property stated verbatim as a design goal. The orchestration layer composes agent outputs depth-by-depth over SQS. (Source: sources/2026-09-30-aws-how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock)
- sources/2026-09-29-dropbox-evolving-our-calendar-assistant-reclaim-to-be-ai-native-with-f60d451f — per-request context/tool scoping + subagent delegation in a calendar assistant. Reclaim's agent platform gives the agent only the calendar information and tools relevant to each request "rather than access to everything at once" (the scoping half of the pattern: shrink the reasoning space per task), and for a more complex request "the agent can give one specific part to a specialized subagent," with internal checklists and review steps keeping the overall request on track. A consumer-product instance of the same intent as the SecOps/observability decompositions below — decompose to keep each agent's tool inventory and context small — though the collaboration protocol between agent and subagent is described only at a high level.
- sources/2026-09-29-cloudflare-adaptive-application-security-for-the-ai-era-how-cloudflare — three-stage autonomous security-operations decomposition. Cloudflare's autonomous SecOps platform splits the investigation into specialized stages: deterministic workflows establish customer + investigation context (trigger history, traffic baselines, enforcement outcomes, network observations); a detection agent searches authorized datasets for anomalies and correlations; specialist agents then review the evidence alongside customer history + threat intelligence to connect isolated events into a campaign. The system recommends mitigations (rate limiting, WAF, DDoS changes) but stops at a human-approval gate — a bounded-authority decomposition where each agent has a narrow role and the final act stays with a person. Co-developed with the Managed Defense SOC team. Security-operations sibling to Slack SPEAR's investigation agents and Cloudflare's own vulnerability-discovery harness.
-
sources/2026-09-25-databricks-from-data-to-dialogue-how-sp-global-energy-made-its-structured-data-estate-conversational — One Genie Agent per dataset group, not per commodity. S&P Global Energy deliberately builds "a fleet of small, sharply scoped Genie Agents rather than a handful of sprawling disconnected AI tools": within LNG alone, separate Assets & Contracts / Cargo / Tenders / Outages / Supply & Demand / Netbacks / Prices agents; Chemicals split across capacity, production, utilization, trade, demand, inventory, and supply–demand balances. The stated rationale is the pattern's exact intent — a single giant agent "degrades answer quality," so each agent carries a narrow tool/context scope and answer accuracy is highest "when an agent covers one specific data domain with clear instructions and example queries." Cross-domain reach is provided above the agents at the MCP layer — a FastMCP proxy composes group Genies into per-commodity endpoints (patterns/mcp-as-centralized-integration-proxy) whose LLM routes or fans-out — rather than by widening any single agent. Best practice, verbatim: "Resist the temptation to build one agent per commodity — or worse, one agent to rule them all. Bring together multiple data domains at the MCP layer instead." A structured-data-estate instance of the pattern (sibling to the Databricks Storex, Slack Spear, and security-review instances below).
-
sources/2025-12-03-databricks-ai-agent-debug-databases — Databricks' systems/storex enables "specialized agents for different domains: one focused on system and database issues, another on client-side traffic patterns, and so on. This decomposition enables each agent to build deep expertise in its area while collaborating with others to deliver a more complete root cause analysis. It also paves the way for integrating AI agents into other parts of our infrastructure, extending beyond databases."
-
sources/2026-04-20-cloudflare-orchestrating-ai-code-review-at-scale — Cloudflare's AI Code Review system is the canonical wiki instance of the pattern applied to code review. Seven specialised sub-reviewers (security, performance, code quality, documentation, release, AGENTS.md, engineering- codex) run in parallel, each with a tightly scoped prompt and an explicit "What NOT to flag" section. Coordinated by a judge-pass coordinator on the top model tier. See the specialisation-dedicated pattern patterns/specialized-agent-decomposition and the orchestration shape patterns/specialized-agent-decomposition. Production scale (first 30 days): 131,246 runs across 5,169 repos; 85.7% prompt-cache hit rate; ~1.2 findings per review.
-
sources/2025-12-01-slack-streamlining-security-investigations-with-agents — canonical security-investigation instance. Slack's Spear applies the pattern at the Expert-agent layer: four specialised security-domain experts (Access — authentication/authorization/perimeter, Cloud — infrastructure/compute/orchestration/networking, Code — source-code + configuration-management analysis, Threat — threat analysis + intelligence). Each Expert owns a distinct toolset / data-source set tied to its domain; the Director broadcasts a question to all four in discovery phase and picks one in trace phase. Distinct architectural sibling structure: Slack composes the peer-Expert layer with a supra-agent (Director) + a meta-agent (Critic) rather than the Databricks peer-collaboration framing — see director-expert-critic-investigation-loop for the three-persona shape. Cost-tier strategy is explicit (knowledge pyramid): Experts on cheap models because leaf work is tool-call-heavy and cognitively shallow; Critic on mid-tier; Director on top-tier. Canonical emergent-behaviour payoff: Critic caught a credential exposure the Expert missed, Director pivoted the investigation.
-
sources/2025-11-17-dropbox-how-dash-uses-context-engineering-for-smarter-ai — Dropbox Dash extracts query construction for its universal search tool into a dedicated search sub-agent. The main planning agent decides when to search; the sub-agent owns the how (user-intent → index-field mapping, query rewriting for semantic matching, typos / synonyms / implicit context). Named rationale: "When a tool demands too much explanation or context to be used effectively, it's often better to turn it into a dedicated agent with a focused prompt." This is the pattern applied to sub-tool complexity, not just domain separation — same shape, different motivation from Storex.
-
sources/2026-01-28-dropbox-knowledge-graphs-mcp-dspy-dash — Josh Clemm's companion talk extends the Dash decomposition with an additional mechanism: a classifier picks the sub-agent for complex agentic queries, each sub-agent having a much narrower tool set. "We use a lot of sub-agents for very complex agentic queries, and have a classifier effectively pick the sub-agent with a much more narrow set of tools." This adds a named routing mechanism to the pattern (the classifier) that the 2025-11-17 post didn't explicitly describe; it also positions specialized-agent-decomposition as one of four named fixes Dash applied to make MCP work at scale (alongside patterns/tool-surface-minimization, knowledge-graph-bundle token compression, and tool-result-local-storage).
-
sources/2026-09-09-databricks-evaluation-first-ai-agents-how-zepto-scales-customer-support — Zepto customer-support stack applies the pattern with an explicit vertical + horizontal split under an orchestrator/router (which can hand off to a human at any point). Vertical agents are intent specialists (WIMO order-tracking/ETA, Missing, Expiry, Returns, Quality, Unable-to-Pay, General fallback); horizontal agents are oversight layers that cut across use cases (Image Deduplication; Item Matching & Image-Manipulation Detection). The decomposition's stated payoff is evaluation tractability: metrics compute per vertical (WIMO intent F1, Expiry OCR accuracy) and per horizontal (fraud precision, image-reuse / manipulation detection), so each piece is measured "in isolation and in combination" — a fourth framing where the split is chosen to make per-agent evaluation possible, not just to shrink tool inventory.
Two framings of the same pattern¶
- Domain-based decomposition (Storex). One agent per domain (DB, client traffic, storage, network); composition layer routes cross-domain questions. Intent: scale tool inventory + prompt specialization across many areas of expertise.
- Sub-tool decomposition (Dash). Extract one specific tool's own internal complexity into a sub-agent, because the tool's explanation otherwise starves the parent's context budget. Intent: protect context budget when a single tool's instruction weight grows.
Both converge on the same mechanism (dedicated prompt + dedicated tool surface + orchestration hand-off) for different reasons. A mature production system often does both — per-domain agents plus, within each, sub-agents for the most complex sub-tasks.
AWS reference-architecture shape (2025-12-11)¶
AWS's conversational-observability Strands deployment adopts this pattern with a three-agent split over Kubernetes troubleshooting:
- Agent Orchestrator — coordinates the troubleshooting workflow across the other two agents. "Coordinates troubleshooting workflows."
- Memory Agent — owns conversation context and historical insights across turns / sessions. "Manages conversation context and historical insights."
- K8s Specialist — narrow-surface diagnostic agent calling EKS MCP Server tools. "Handles Kubernetes diagnostics."
The decomposition mirrors Storex (per-storage-layer specialists), Dash (classifier-routed sub-agents), and Cloudflare Agent Lee (domain-per-team agents): same pattern, same rationale — keeping each agent's tool inventory small enough for reliable selection and small enough to fit in context. Operational-ops instance of the pattern, same shape as Storex's storage-incident instance. (Source: sources/2025-12-11-aws-architecting-conversational-observability-for-cloud-applications)
Verification-gated inner-loop variant (DS-STAR, 2025-11-06)¶
Google Research's DS-STAR data-science agent is the canonical wiki instance of specialised-agent decomposition organised by role in a refinement loop, not by subject-matter domain:
- Data File Analyzer — writes + runs a file-summarisation script.
- Planner — emits the high-level plan.
- Coder — turns the plan into executable code, runs it.
- Verifier — LLM judge scoring plan sufficiency against intermediate results.
- Router — on reject, decides add-step vs fix-step.
The agents specialise not because they own different domains but because they own different roles in the plan → implement → verify → refine loop; the Verifier gates each cycle, and the Router's add-or-fix decision is the refinement primitive. Full pattern is patterns/specialized-agent-decomposition; loop-level concept is concepts/context-engineering.
Ablations quantify the decomposition's value: removing the Data File Analyzer collapses DABStep hard-task accuracy 45.2 % → 26.98 %; removing the Router (forcing extend-only) degrades both easy and hard tasks — "it is more effective to correct mistakes in a plan than to keep adding potentially flawed steps" (Source: sources/2025-11-06-google-ds-star-versatile-data-science-agent).
Adds a third framing to this pattern's taxonomy alongside the Storex (domain-based) and Dash (sub-tool) framings:
- Role-in-the-refinement-loop decomposition (DS-STAR). One agent per loop role (context, plan, implement, judge, route); coordination is the loop itself. Intent: make verification and revision first-class, isolate ablation-testable primitives.
Offline-context-generation framing (Meta, 2026-04-06)¶
Meta's AI Pre-Compute Engine is the canonical wiki instance of the pattern applied offline — a one-session orchestration of 50+ specialised agents that reads a 4-repo / 4,100-file config-as-code data pipeline and emits 59 context files plus a dependency graph. Nine named roles:
- Explorers (2) — map the codebase.
- Module analysts (11) — apply the five-questions framework per module.
- Writers (2) — synthesise the 59 compass-shaped context files.
- Critics (10+ across 3 rounds) — independent quality review.
- Fixers (4) — apply corrections.
- Upgraders (8) — refine the orchestration / routing layer.
- Prompt testers (3) — validate 55+ queries × 5 personas.
- Gap-fillers (4) — cover remaining directories.
- Final critics (3) — integration tests.
This is the fourth framing: the agents specialise by pipeline stage in an offline knowledge-extraction flow. The Storex / Dash / DS-STAR framings are runtime decompositions; Meta's is a one-shot offline orchestration whose output is a durable artifact (the 59 files) that downstream runtime agents consume. Delivers measurable outcomes: critic quality 3.65 → 4.20 / 5.0 across 3 rounds, zero hallucinated file paths, ~40 % fewer tool calls per task on a six-task preliminary eval (Source: sources/2026-04-06-meta-how-meta-used-ai-to-map-tribal-knowledge-in-large-scale-data-pipelines).
Skill-over-shared-tools framing (Meta Capacity Efficiency, 2026-04-16)¶
Meta's Capacity Efficiency Platform is the canonical wiki instance of the pattern applied as skill-based composition over a shared MCP tool layer. The same five tools (profiling · experiments · config history · code search · documentation) serve every specialist agent; the agents differ only in their skill bundle (named encoded domain expertise).
At least seven specialist agents share the platform:
- AI Regression Solver — defense, runs on top of FBDetect, produces fix-forward PRs (ai-generated-fix-forward-pr).
- Opportunity Resolver — offense, turns proactive efficiency opportunities into candidate fixes in the engineer's editor (opportunity-to-pr-ai-pipeline).
- Conversational efficiency assistants.
- Capacity-planning agents.
- Personalised opportunity recommendations.
- Guided investigation workflows.
- AI-assisted validation.
Meta's explicit claim: "each new capability requires few to no new data integrations since they can just compose existing tools with new skills." This is a fifth framing to this pattern's taxonomy: the agents specialise by skill bundle, not by subject-matter domain (Storex) / sub-tool complexity (Dash) / refinement-loop role (DS-STAR) / offline-pipeline stage (Pre-Compute Engine). Differentiator: the tool layer is shared across specialists, and new agents cost only new skill-authoring, not new data integration (Source: sources/2026-04-16-meta-capacity-efficiency-at-meta-how-unified-ai-agents-optimize-performance-at-hyperscale).
See mcp-tools-plus-skills-unified-platform for the full architectural pattern.
Regulated-financial-services framing (IBM + AWS KYC, 2026-04-23)¶
The IBM + AWS KYC architecture is the canonical wiki instance of specialised-agent decomposition applied to regulated compliance workflows on Bedrock AgentCore. One KYC Orchestration Supervisor + five domain sub-agents:
- Identity Verification — watchlist / sanctions APIs, name- variation NLP.
- Document Analysis — OCR, multi-language, watermark / security-feature forgery detection.
- Fraud Detection — behavioural analysis, same-IP / same- device collision detection, semantic-similarity search over historical fraud, dynamic risk scoring with explainable assessments.
- Compliance & Risk — jurisdiction-specific regulatory interpretation (BSA, USA PATRIOT Act, AMLD, MAS, FATF), attestation generation with audit trails.
- Customer Experience — real-time friction-point detection, abandonment-reduction recommendations.
Supervisor does no compliance work itself; it dynamically constructs a parallel-or-sequential execution plan per case based on document types, geography, risk indicators, and historical patterns. This is a sixth framing in this pattern's taxonomy: regulated-compliance decomposition — agents specialise by regulatory sub-domain, each with a narrowly-scoped OpenAPI tool surface enforced at runtime by systems/agentcore-identity + systems/agentcore-gateway. The decomposition is the key to the sub-5-minute latency target (patterns/specialized-agent-decomposition) and to the composed confidence-tiered routing (>95 / 75-95 / <75) that the Supervisor applies to sub-agent outputs.
Full pattern: supervisor-subagent-kyc-orchestration. (Source: sources/2026-04-23-aws-modernizing-kyc-with-aws-serverless-solutions-and-agentic-ai.)
-
sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore — a four-agent inventory pipeline (Supervisor + Preprocessing + Forecasting + Reporting) implemented as the Agents-as-Tools variant: the Supervisor is one Strands agent whose tool list is the three specialists wrapped as
@tools, orchestrated inside its tool-use loop (max_node_executions=10) rather than an explicit graph. Each specialist gets its own context window and tool subset; results return as compressed CLUES_FORMAT envelopes. Pairs with patterns/deterministic-tool-vs-llm-judgment (what inside each agent is a tool vs reasoning) and patterns/data-contract (typed JSON between agents). Rationale for decomposition over a monolith: a single prompt doing all steps would exceed context limits, be untestable at the component level, and fail catastrophically on any single-step error. -
sources/2026-09-24-databricks-how-i-built-agent-based-security-reviews-on-databricks — security-review-intake instance. A Databricks security leader built the review layer as seven focused agents — Intake, Risk Assessment, Requirements, Specialized Review (e.g. browser-extension threat modeling, vendor assessment), Validation, Workflow, Learning — orchestrated as scheduled Lakeflow Jobs behind a conversational intake app. The decomposition rationale is stated verbatim: the author "deliberately avoided building a single agent with broad authority to act as a security reviewer" because bounded responsibility keeps "its behavior ... inspectable and testable, and a change to one does not silently affect another." Two structural pairings distinguish this instance: (1) a conservative-default risk agent that "defaults to a higher tier when the picture is incomplete" — fail-closed on the security invariant baked into one specialist; and (2) a Learning agent that compares reviewer edits to original output to improve prompts + standards (patterns/human-calibrated-llm-labeling) without mutating prod behavior. Model tiering runs underneath the decomposition rather than per agent: Haiku (classification) → Sonnet (most review) → Opus (heaviest reasoning), the patterns/cheap-approximator-with-expensive-fallback ladder. This is the same shape as Slack Spear (security-domain experts) and Cloudflare AI Code Review (seven scoped sub-reviewers), applied to security intake + review rather than investigation or code review, with human authority preserved for high-risk / novel / ambiguous cases.
-
sources/2026-09-24-google-automating-coherent-long-form-video-generation — creative-production instance (Co-Director). Google Research decomposes long-form video generation into a hierarchy: an Orchestrator Agent (multi-armed-bandit strategy selection) → a Pre-Production Agent (storyboard) → a Production Agent with three specialized sub-agents — Keyframe (anchors character/scene), Video (adds motion), Audio (voiceover + score). Rationale matches the pattern's thesis: independent handcrafted prompting causes drift and cascading failures that are hard to attribute (the credit-assignment problem), so each agent carries a small scope under a single unified vision injected top-down into its system prompt. An MLLM Judge (concepts/llm-as-judge) closes a factored-reward loop back to the orchestrator. Sibling frameworks CANVAS (world-state memory) and VQQA (closed-loop refinement) are themselves multi-agent. First creative/generative-media instance of this pattern (vs. investigation, code review, KYC, forecasting).
Related¶
- patterns/specialized-agent-decomposition — the enabling framework.
- patterns/snapshot-replay-agent-evaluation — per-agent eval.
- systems/storex
- systems/strands-agents-sdk
- concepts/agentic-development-loop
- patterns/specialized-agent-decomposition — role-in-the-refinement-loop variant of this pattern.
- systems/ds-star — canonical instance of role-based decomposition with inner-loop verification.
- concepts/context-engineering — the loop-level discipline.
- systems/meta-ai-precompute-engine — offline-context-generation variant.
- precomputed-agent-context-files — the containing offline pattern.
- supervisor-subagent-kyc-orchestration — regulated-compliance framing.
- systems/bedrock-agentcore — runtime substrate for the regulated-compliance framing.
Merged aliases¶
domain-decomposed-research-workflowmulti-agent-review-coordinatorparallel-narrow-agents-over-exhaustiveparallel-subagent-execution-for-latencyphase-scoped-specialized-agentsplanner-coder-verifier-router-loopspecialized-reviewer-agentsswarm-of-discovery-agents-for-context-prebuild-adversarial-review-subagentclassifier-based-smart-routingcoordinator-sub-reviewer-orchestrationevidence-fan-out-then-synthesizemulti-agent-debate-evaluationmulti-agent-supervisor-routingtask-level-routing-via-meta-harnesstool-decoupled-agent-framework