Skip to content

CoreBreak proves agent guardrails need to live outside the agent

Summary

A Redpanda governance-thesis post (2026-08-31, unsigned) built around CoreBreak — a structural agent-harness vulnerability class disclosed by researchers Aviyam Ivgi and Hedi Ingber at Black Hat 2026 that ran through three of the largest agent stacks: AWS Bedrock AgentCore with the Strands SDK, Google's Agent Development Kit (ADK), and Vercel's AI SDK harness. All three vendors patched it under emergency CVEs (AWS CVE-2026-18830, CVSS 8.6; Google CVE-2026-18236, CVSS 9.3 critical; Vercel CVE-2026-64650 + CVE-2026-64651). The post's thesis: every industry guardrail (system prompts, content filters, refusal training) is wrapped around the model — so bypass the model and you bypass all of them. The structural fix all three vendors converged on is out-of-band policy enforcement (OBPE): bind tool execution to state the attacker cannot author, and reject externally-supplied tool calls and approvals. Redpanda argues its Agentic Data Plane shipped that OBPE boundary before the breach, and that a boundary is stronger than a provenance check because it removes the prize for defeating provenance rather than merely detecting a forged call.

Key takeaways

  1. CoreBreak is a single structural flaw across three major agent stacks. The guardrails all sit around the model; skip the model and the guardrails never fire. Verbatim: "If you bypass the model, you bypass all of them." (Source: this post.) → concepts/prompt-injection

  2. Attack #1 — forged tool call in message history. "Feed the harness a message history whose latest message already contains a tool call, and the tool runs with attacker-chosen arguments while the model and every guardrail wrapped around it never fires." This is a confused deputy executed by fabricating the model's own output rather than by prompt injection of the model's input.

  3. Attack #2 — forged approval event. "Forge an 'approved' confirmation event in that same history, and the human-in-the-loop step approves an action no human saw. The approval mechanism itself became the attack vector." → forged-approval-event

  4. The emergency fixes all converged on OBPE. "Bind tool execution to state the attacker cannot author, and reject externally supplied tool calls and approvals." Each vendor "had to rebuild a piece of it under fire, one patch at a time." The vendors' patches add provenance checks — verify a tool call came from a model turn, reject the rest. → out-of-band-agent-enforcement

  5. A boundary is not the same as a provenance patch — it removes the prize for defeating it. "A boundary doesn't replace that check; it removes the prize for defeating it. A forged call that gets past provenance still meets the data owner's policy ceiling, so it can do at most what a legitimate call could, not whatever the execution role's full reach would allow." For approvals the case is stronger still: an OBPE approval "never reads from agent-supplied history at all, so the forged-confirmation trick has nothing to forge." → boundary-removes-prize-for-defeating-provenance

  6. OBPE beats prompt-only and in-band-reviewer guardrails empirically. Redpanda's OBPE research paper ran 3,621 trials across four models (Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5, GPT-5 mini) against a prompt-injection / forbidden-action attack class:

  7. Prompt-rule-only agents failed 57.6% of trials (leaking protected data, revealing protected values, or completing forbidden backend effects).
  8. The same agents behind the OBPE boundary failed 0.2%.
  9. An in-band reviewer model (reads agent output, blocks unsafe answers) "still failed 58.6% of the time" on Sonnet — "It can refuse an answer. It cannot un-read a record the agent has already absorbed, and it cannot undo a write that already landed." The boundary failed 0.0% of the same trials, and did more useful work: 66.2% task fulfillment vs the reviewer's 58.7%, and 63.0% safe-useful completion vs 19.7% (a ~3× improvement) while the reviewer "made 930 extra model calls."

  10. A policy decision is deterministic and cheap; a model guard is not. "A model guard is an inference call on every step, so its cost and latency scale with its capability. You cannot buy your way from probabilistic to deterministic. A policy decision is a typed evaluation against a written rule, and it gives you the same answer in microseconds with no tokens spent."

  11. Credential-theft path: sandbox egress is the real prize. CoreBreak's injected instructions steered managed browser + code-interpreter tools into "reading the execution role's cloud credentials and exfiltrating them." "No policy engine changes a sandbox's networking, ours included." The lesson: a typed, mediated tool grant keeps the credential behind the boundary (token vault); "a general compute environment with ambient cloud credentials puts it one HTTP request away from any injected instruction." → reinforces patterns/agent-sandbox-with-gateway-only-egress + four-component-agent-production-stack.

  12. The self-approval framing. "No company lets employees approve their own access requests from a document they wrote themselves. Agents shouldn't either." The question to ask of every agent stack: "whether your agent is governing itself. CoreBreak is what happens when the answer is 'yes.'"

Systems / concepts / patterns extracted

Operational numbers

  • 3,621 trials, four models (Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5, GPT-5 mini).
  • Prompt-rule-only failure rate: 57.6%; OBPE boundary: 0.2%.
  • In-band reviewer (Sonnet) failure rate: 58.6%; boundary on the same trials: 0.0%.
  • Task fulfillment: boundary 66.2% vs reviewer 58.7%.
  • Safe-useful completion: boundary 63.0% vs reviewer 19.7% (~3×).
  • Reviewer overhead: 930 extra model calls; boundary spent none.
  • CVEs: CVSS 8.6 (AWS), 9.3 critical (Google), two Vercel CVEs.
  • Policy-decision latency: "the same answer in microseconds with no tokens spent."

Caveats

  • Vendor-thesis post, not a mechanism deep-dive. The architecture of the OBPE boundary (identity binding, request-digest-tied approvals, immutable held-call resume) is asserted at claim altitude; the mechanism is deferred to the research paper, the ACM CAIS SAO workshop paper, and the O'Reilly Radar essay.
  • The 3,621-trial benchmark tests a sibling attack class, not CoreBreak itself: "not forged tool calls, but prompt-injected agents coaxed into leaking protected data and taking forbidden actions." The post argues the same design answer covers both; the CoreBreak forged-call / forged-approval class is argued structurally, not benchmarked here.
  • Self-promotional frame. The post is explicitly "we built the OBPE boundary before the breach" positioning for the Agentic Data Plane; CVE facts and the researchers' hardening list are the load-bearing third-party corroboration.
  • CoreBreak researcher hardening list quoted, not the full paper. The post paraphrases the researchers' recommendations (reject caller-authored tool calls, authorize at execution time, reduce inherited authority, bind each invocation to the model event + auth state that produced it) as "a specification for enforcement outside the agent."
  • Tier-3 source. Redpanda blog; passes scope on genuine agent-security architecture content (attack mechanics, OBPE boundary vs provenance distinction, benchmark numbers) well over the 20% threshold.

Source

Last updated · 766 distilled / 2,225 read