Skip to content

SYSTEM Cited by 3 sources

Out-of-Band Policy Engine (OBPE)

Definition

The Out-of-Band Policy Engine (OBPE) is the Govern pillar of the Redpanda Agentic Data Plane: a policy engine that enforces what an AI agent may see and return at the MCP boundary, server-side, in the path of every call — rather than inside the agent. For any tool on any MCP server, an operator declares field-level rules: mask a field, drop it, or clamp an argument to a legal range. "The agent never sees the rule or the unmasked value, and has no surface to negotiate with."

Introduced 2026-08-03 (Source: sources/2026-08-03-redpanda-out-of-band-policy-engine-governance-ai-agents-cant-ignore). The name encodes the design tenet: out-of-band enforcement — the policy lives where the agent cannot read, modify, or route around it.

What it enforces

Verbatim: "For any tool on any MCP server, you declare what the agent may see: mask a field, drop it, clamp an argument to a legal range. It applies server-side, in the path, on every call."

  • Mask — replace a field's value (e.g. redact PII) before it reaches the agent.
  • Drop — remove a field from the tool response entirely.
  • Clamp — bound a tool argument to a legal range before the call executes.

Schema rewrite so masking doesn't break the contract

A subtle but load-bearing detail: "When masking changes a field's type, OBPE rewrites the tool schema to match, so the agent gets a coherent contract instead of a type error." The design principle stated: "Governance that makes agents dumber doesn't survive contact with the people deploying them." Enforcement must be invisible and leave the agent a well-typed tool contract.

Policy attaches to the resource, not the agent

The organizational-scale property is per-tenant-policy-store:

"Policy attaches to the MCP server rather than to individual agents, which is what makes this work at organizational scale. The team that owns a data source sets the ceiling once, for every agent that will ever connect (and can vary it by user, group, or agent)… The data owner doesn't have to review every agent the company subsequently builds."

OBPE also governs MCP servers Redpanda didn't write — "self-managed, remote, third-party, the one your teammate vibe-coded on Tuesday."

The failure mode it fixes (not an attacker)

OBPE targets the confused-deputy shape, not a breach: "Your triage agent runs on your credential… It needs one Jira board, but the credential opens all fifty. A customer pasted an API key into a ticket description… Your agent reads the ticket. Someone asks it an innocent-sounding question. It reads the key back out loud. Nothing was hacked, and every permission check passed." The authorization was the wrong size and the data crossed a boundary it should never have — "you can't patch that, and you can't align your way out of it." OBPE makes the boundary a server-side field-level rule instead of a hope about the agent's behaviour.

What ships with it

  • Attribute-based access control (ABAC) — fine-grained, attribute-driven permissions across agents, tools, and connections; the substrate role/group models sit on. "Tag data internal-only, and agents serving external users can't reach it regardless of the credential."
  • Bring-your-own-agent transcripts — LangChain / CrewAI / ADK / hand-rolled agents get the same full-fidelity record; no rewrite.
  • AWS Bedrock Guardrails integration — "Bedrock providers today. OBPE's field-level enforcement is provider-agnostic."
  • Cost attribution by team / department / project; scheduled agent runs (cron-like, timezone-aware); a CLI for agents, gateways, and MCP servers.

Relationship to sibling systems / patterns

  • Agent Network View — the See pillar; OBPE's authorization-denial and guardrail-violation decisions are the overlays rendered on that graph.
  • Agent kill switch — the Stop pillar (coming soon). OBPE constrains each call; the kill switch cuts the whole agent off.
  • mcp-boundary-field-level-enforcement — the general pattern OBPE instantiates (enforce mask/drop/clamp + schema-rewrite at the protocol boundary, server-side).
  • patterns/central-proxy-choke-point — OBPE works because the Agentic Data Plane is a single choke point in the path of every call.
  • Databricks UC Service Policies — the closest cross-vendor sibling: per-call policy attached to the resource (MCP server / UC function), not the agent. OBPE is the streaming-vendor instance; Service Policies the lakehouse-vendor instance.

Caveats

  • Mechanism depth partial. How the schema-rewrite is computed, how policies compose when user/group/agent scopes overlap, and the in-path latency cost of server-side enforcement are not disclosed.
  • In-path coverage required. OBPE can only enforce on calls that traverse the MCP boundary it governs; egress outside the plane is the whole-path gap.

CoreBreak validation (2026-08-31): boundary > provenance patch

The CoreBreak disclosure at Black Hat 2026 (Source: sources/2026-08-31-redpanda-corebreak-proves-agent-guardrails-need-to-live-outside-the-a-6ef5e77e) is Redpanda's strongest external argument for OBPE. Three major agent harnesses (AWS Bedrock AgentCore + Strands, Google ADK, Vercel AI SDK) were found to execute forged tool calls and forged approvals injected into the message history, bypassing every model-wrapped guardrail; all three patched under emergency CVEs.

  • OBPE is not the same as the CVE patches. The vendor patches add provenance checks (verify a call came from a model turn, reject the rest). OBPE removes the prize for defeating that check: "A forged call that gets past provenance still meets the data owner's policy ceiling, so it can do at most what a legitimate call could, not whatever the execution role's full reach would allow."
  • Approvals bound to unforgeable state. "Each approval ties to the exact request digest, the deciding principal, and the policy version, and resumes only the immutable held call." Because an OBPE approval "never reads from agent-supplied history at all," the forged-confirmation trick "has nothing to forge."
  • Measured. In Redpanda's 3,621-trial OBPE research benchmark across four models, prompt-rule-only agents failed 57.6%; the same agents behind the OBPE boundary failed 0.2%. An in-band reviewer model failed 58.6% (Sonnet) while the boundary failed 0.0% and did ~3× more safe-useful work (63.0% vs 19.7%). "A policy decision is a typed evaluation against a written rule … the same answer in microseconds with no tokens spent."

The "kill switch is downstream" framing (2026-09-28)

The 2026-09-28 "Your AI kill switch is in the wrong place" post (Source: sources/2026-09-28-redpanda-your-ai-kill-switch-is-in-the-wrong-place) is the policy-audience restatement of OBPE's thesis, occasioned by the AI Kill Switch Act and the Hugging Face / OpenAI sandbox-escape incident. It names the reusable framing behind OBPE's whole reason to exist — the business-context gap: sandboxes, network allowlists, and scoped API keys control what an agent can reach, not whether an action is right for the task, so "an agent deleting an old test environment and one deleting production… both make the same kind of API call with valid permissions." Only the boundary — judging who asked, what task, why — can tell them apart. Two OBPE correctness properties are restated as the post's core: determinism ("the same request in the same context must get the same decision every time" — a typed evaluation, not the agent's own understanding) and immutable audit ("each decision lands in an audit record the agent can't alter"). The post also fixes the ordering between OBPE and the Stop pillar: the boundary decides allow / block / hold-for-human before the call and re-checks the result after, and the kill switch "sits downstream of all of it… the last resort," fired by the boundary's verdicts rather than substituting for them.

Seen in

  • sources/2026-08-03-redpanda-out-of-band-policy-engine-governance-ai-agents-cant-ignore — introduction (2026-08-03) as the Govern pillar; mask/drop/clamp + tool-schema-rewrite at the MCP boundary, policy attached to the MCP server, ships with ABAC + BYO-agent transcripts + Bedrock Guardrails + cost attribution.
  • sources/2026-08-31-redpanda-corebreak-proves-agent-guardrails-need-to-live-outside-the-a-6ef5e77e — CoreBreak validation (2026-08-31): OBPE as the structural answer the CVE patches converged on, the policy-ceiling-beats-provenance argument, unforgeable request-digest-tied approvals, and the 3,621-trial benchmark (57.6% → 0.2%).
  • sources/2026-09-28-redpanda-your-ai-kill-switch-is-in-the-wrong-place — policy-audience restatement (2026-09-28): the business-context gap (access ≠ permission), the model-vs-agent enforcement-layer argument, the three-database example (identical API calls, different business context), OBPE's allow/block/hold-for-human boundary with deterministic decisions + immutable audit, and the reconciliation that the kill switch is downstream of the boundary, not the boundary itself. No new mechanism over the 2026-08-03 post; adds the Hugging Face / OpenAI incident and the AI Kill Switch Act as external framing.

Source

Last updated · 766 distilled / 2,225 read