The out-of-band policy engine: governance AI agents can't ignore¶
Summary¶
A Redpanda blog post (2026-08-03, founder-voice) that states the load-bearing principle behind Redpanda's Agentic Data Plane governance layer and announces its latest release along three verbs — See, Govern, and (coming soon) Stop. The thesis: governance must be enforced through channels the agent cannot access, modify, or circumvent — "if the agent can see the policy, you don't have governance. You simply gave it a suggestion." The release delivers three concrete pieces: (1) Agent Network View — a live graph of the agent estate derived from traffic, not from a registry someone remembered to update (the See answer to agent sprawl); (2) the Out-of-Band Policy Engine (OBPE) — field-level masking / dropping / argument-clamping enforced server-side at the MCP boundary, with policy attached to the MCP server (the resource) rather than to individual agents (per-tenant-policy-store); and (3) the agent kill switch (Stop, coming soon) — surgical instance-level cutoff. The architecturally novel claim is that agent governance is a data-infrastructure problem, not a dashboard problem: every input/output/tool-call is written to a log — a Kafka-compatible topic in the customer's own account, not a dashboard, so a program (a continuous grader) can subscribe to it and its verdict stream can drive the kill switch on evidence rather than on a human eventually noticing.
Key takeaways¶
-
The out-of-band principle (thesis). "Governance must be enforced through channels the agent cannot access, modify, or circumvent." Every in-agent control — rules in the system prompt, a guardrail model on the output, an in-process permission check — shares one fatal property: the agent can see it, so a malicious/hijacked agent can alter, disable, or delete its own guardrails and the audit trail of what it just did. "Out-of-band means out of the agent's reach. The agent doesn't get a vote. It doesn't get to know there even was a vote." Canonical statement of out-of-band-agent-enforcement (Source: this post).
-
Why fixing the agent doesn't work. "Your security model is exactly as strong as a language model's willingness to remember and obey words under every condition, including adversarial ones. Guard models don't escape this; they just add a floor to it, so now there's a second non-deterministic, injectable system supervising the first. It's just LLMs all the way down." An in-process check "only fires if the agent calls it instead of the API underneath. Agents write and run code. That's an expectation, not a guarantee."
-
See — Agent Network View. A live graph of every agent, the models it calls, the MCP servers it connects to, the policies applied, and the tools it actually invoked — with runtime overlays for guardrail violations and authorization denials, plus a KPI strip (agents, conversations, tool calls, tokens, spend). "Because the plane is in the path of every model call and every tool call, the topology isn't reconstructed from a survey or a registry somebody remembered to update. It's derived from traffic, so if an agent is running, it's on the graph." Requires no cooperation from the agents and no knowledge that they're watched — the See answer to agent sprawl / shadow AI. See systems/redpanda-agent-network-view.
-
Govern — Out-of-Band Policy Engine (OBPE). "For any tool on any MCP server, you declare what the agent may see: mask a field, drop it, clamp an argument to a legal range. It applies server-side, in the path, on every call. The agent never sees the rule or the unmasked value, and has no surface to negotiate with." When masking changes a field's type, OBPE rewrites the tool schema to match, so the agent gets a coherent contract instead of a type error. Enforcement is at the MCP boundary rather than inside the agent. See systems/out-of-band-policy-engine.
-
Policy attaches to the resource, not the agent. "Policy attaches to the MCP server rather than to individual agents, which is what makes this work at organizational scale. The team that owns a data source sets the ceiling once, for every agent that will ever connect (and can vary it by user, group, or agent)… The data owner doesn't have to review every agent the company subsequently builds." This is per-tenant-policy-store — the same inversion Databricks calls governance travels with resources, not frameworks and the same central-choke-point stance. It also governs MCP servers Redpanda didn't write: "self-managed, remote, third-party, the one your teammate vibe-coded on Tuesday."
-
The failure mode OBPE fixes — and it isn't an attacker. "Your triage agent runs on your credential… It needs one Jira board, but the credential opens all fifty. A customer pasted an API key into a ticket description… Your agent reads the ticket. Someone asks it an innocent-sounding question. It reads the key back out loud. Nothing was hacked, and every permission check passed." The authorization was the wrong size and the data crossed a boundary it never should have — the confused-deputy shape. "You can't patch that, and you can't align your way out of it."
-
Governance is a data-infrastructure problem, not a dashboard problem (the load-bearing architecture claim). Every input, output, and tool call is written to a transcript — "Not sampled or summarized. And critically, not written to a dashboard. Instead, it's written to a log: a Kafka-compatible topic you can consume, in your own account, queryable with standard Postgres semantics." "A dashboard is something a human reads after the fact… A log is something a program subscribes to." This is log-not-dashboard — it lets you "hang a grader off it that evaluates agent behavior continuously, in real time," and that grader's verdicts "become their own log, and a verdict stream is exactly what a kill switch needs upstream of it."
-
Stop — the agent kill switch (coming soon). "Cut a misbehaving agent off from the gateway and every tool it can reach, in one action." Redpanda explicitly ships it after the record: "A switch is only as good as whatever tells it to fire, and nothing can tell it to fire without a complete record of what your agent actually did, sitting somewhere the agent cannot edit. Build the switch first and you've built a button nobody knows when to press. That must be the order." See concepts/ai-agent-guardrails.
-
Four questions to tell governance from security theatre. (1) Is it out of band? — "if the policy lives anywhere the agent can read, modify, or route around, nothing else on this list matters." (2) Does it cover the whole path? — model call + tool call + data access
-
compute + the record of all of it; "any segment left uncovered is where the sprawl leaks out" (an AI gateway covers one segment, an observability vendor a different one, a prompt firewall part of one) → whole-path-governance. (3) Does it work with what you already have? — any model / framework / client / MCP server via open standards (MCP, A2A, OIDC, OAuth, OTel, OCSF, Kafka); "Governance you only get for the agents you bought from one vendor… is just a walled garden." (4) Does it run where your data already is? — BYOC into your own cloud account; "full-fidelity transcripts are the single most sensitive artifact this entire system produces… In a vendor's SaaS, it's a second copy of your crown jewels."
-
Mechanical vendor severance mirrors the kill switch, one layer down. "Our control plane never sees your data. It authors the desired state and nothing more, and it holds no standing credentials in your account. Management traffic is outbound from your network only. And the whole thing is built so that a single firewall rule severs all of our access while your data plane keeps serving. You should be able to fire your vendor mechanically, not just contractually." This is mechanical-vendor-severance — the control/data-plane separation tenet applied to the vendor relationship itself.
Systems / concepts / patterns extracted¶
Systems: Agentic Data Plane (the governance layer this release extends), Agent Network View (NEW — the See pillar), Out-of-Band Policy Engine (OBPE) (NEW — the Govern pillar), MCP (the enforcement boundary), AWS Bedrock Guardrails (integrated provider), Kafka (the transcript-log substrate — the "Kafka-compatible topic"), Redpanda SQL (the "queryable with standard Postgres semantics" surface over the log), Redpanda (the streaming substrate).
Concepts: out-of-band-agent-enforcement (canonical principle statement), log-not-dashboard (NEW — governance record as subscribable log, not read-after-the-fact dashboard), whole-path-governance (NEW — cover model+tool+data+compute+ record, not one segment), coding-agent-sprawl (the problem the See pillar solves), concepts/ai-agent-guardrails (the Stop pillar), concepts/governed-agent-data-access, concepts/attribute-based-access-control (ABAC shipped in OBPE), concepts/control-plane-data-plane-separation, concepts/confused-deputy-problem, concepts/data-residency.
Patterns: per-tenant-policy-store (NEW — data owner sets the ceiling once at the MCP server for every agent), mcp-boundary-field-level-enforcement (NEW — mask/drop/clamp + schema-rewrite server-side at the MCP boundary), verdict-stream-upstream-of-killswitch (NEW — a continuous grader off the transcript log emits verdicts that trip the kill switch), mechanical-vendor-severance (NEW — one firewall rule severs vendor access while the data plane keeps serving), patterns/central-proxy-choke-point (the general out-of-band form).
Operational specifics / features shipped¶
Shipped in this release (the See + Govern pillars):
- Agent Network View — live agent-estate graph derived from traffic; runtime overlays for guardrail violations + authorization denials; KPI strip (agents, conversations, tool calls, tokens, spend).
- Out-of-Band Policy Engine (OBPE) — per-tool, per-MCP-server field-level enforcement: mask / drop / clamp-to-legal-range; server-side, in-path, on every call; tool-schema rewrite on masks that change a field's type; policy varies by user / group / agent.
- Attribute-based access control (ABAC) — attribute-driven permissions across agents, tools, connections; e.g. tag data internal-only and external-serving agents can't reach it regardless of credential. See concepts/attribute-based-access-control.
- Bring-your-own-agent transcripts — LangChain / CrewAI / ADK / hand-rolled agents get the same full-fidelity record; no rewrite.
- AWS Bedrock Guardrails integration — "Bedrock providers today. OBPE's field-level enforcement is provider-agnostic."
- Cost attribution by team / department / project — spend already visible per agent; tagging rolls it up to the org chart.
- Scheduled agent runs — cron-like with timezone control.
- CLI — manage agents, gateways, MCP servers from the terminal.
Early-preview / coming-soon:
- Agent kill switch — surgical, instance-level cutoff from the gateway and all tools.
- Transcript + audit export — managed streaming of transcripts + OCSF audit records to DataDog or your own SIEM.
- Audit log in OCSF format — every governed action in the SIEM's schema ("Format now, managed export soon").
- Agent edit controls + versioning — gate who can change an agent's prompt, with full config history.
- More managed MCP connectors.
Deployment / posture:
- BYOC into your own cloud account, on your own network — "Bring-your-own-cloud (BYOC) is not a premium tier; it's the only way this ships. There's no vendor-hosted version." Ties to concepts/data-residency and the "where your data already is" question. "Hundreds of BYOC clusters on AWS, Azure, and GCP."
- Control plane holds no standing credentials; management traffic is outbound-only; a single firewall rule severs all vendor access while the data plane keeps serving.
Caveats¶
- Product-launch post with a genuine architecture core. This is a See/Govern/Stop release announcement in founder-voice, included per AGENTS.md borderline rule because the out-of-band principle, the MCP-boundary field-level enforcement design, the log-not-dashboard substrate claim, and the resource-attached-policy inversion are real, reusable system-design content (well over 20% of the body).
- Kill switch is coming-soon, not shipped. The Stop pillar and the transcript/OCSF export are explicitly early-preview; only See + govern (Agent Network View + OBPE) are stated as shipped.
- No quantitative numbers. No latency figures for in-path OBPE enforcement, no throughput for the transcript log, no scale figures for Agent Network View beyond "hundreds of BYOC clusters."
- Founder-voice / marketing register. Gallego-style prose ("It's just LLMs all the way down", "the one your teammate vibe-coded on Tuesday") with heavy links to companion posts (four-pillars, the kill-switch/Redpanda-SQL post). The load-bearing claims are the out-of-band principle, MCP-boundary enforcement, and log-not-dashboard.
- OBPE mechanism depth partial. How the schema-rewrite is computed, how policies compose when user/group/agent scopes overlap, and the in-path latency cost are not disclosed.
Source¶
- Original: https://www.redpanda.com/blog/agentic-ai-needs-out-of-band-governance
- Raw markdown:
raw/redpanda/2026-08-03-out-of-band-policy-engine-governance-ai-agents-cant-ignore-1e570767.md
Related¶
- systems/redpanda-agentic-data-plane — the governance layer this release extends (See / Govern / Stop).
- systems/redpanda-agent-network-view — the See pillar; traffic-derived agent-estate graph.
- systems/out-of-band-policy-engine — the Govern pillar; MCP-boundary field-level enforcement.
- out-of-band-agent-enforcement — the canonical principle this post states.
- log-not-dashboard — governance record as a subscribable log, not a dashboard.
- concepts/ai-agent-guardrails — the Stop pillar, driven by the verdict stream.
- per-tenant-policy-store — data owner sets the ceiling once at the MCP server.
- concepts/governed-agent-data-access — Gallego's two-axis governance framing this fits under.
- companies/redpanda — company page.