Detect and send production issues straight to your agent¶
Summary¶
Cloudflare launched Issues (open beta), built-in error monitoring for
Workers that closes the manual gap between
"production is failing" and "an agent is investigating." Coding agents can
already query telemetry, navigate a repo, change code, write tests, and open a
PR — what was still manual was the structured handoff: recognizing that
repeated failures share one root cause, gathering the relevant logs/traces, and
delivering that context to an agent. Issues does three things: (1) groups
repeated exceptions, 5xx responses, and error logs into a single issue; (2)
packages the error, source-mapped stack trace, leading/trailing logs and
traces, Worker version, and application context; and (3) hands it off —
via Automations that fire on an occurrence threshold or return-after-quiet — to
a configured coding agent (Claude Code, Cursor, Devin), a generic webhook, or a
chat/incident tool. Detection is built into the Workers runtime — no SDK,
no application wrapper, one line of wrangler.jsonc config. Cloudflare
dogfooded it on Workflows and, within a day of
turning it on, surfaced and fixed two production bugs by routing the issues into
Cloudflare OS.
Key takeaways¶
- The unit of observability shifts from "raw telemetry" to "the issue." Every failed request has a different request ID but the same bug; Issues groups them and shows first-seen, occurrence count, and whether it is trending up. This is the error-tracking / exception-grouping shape (à la Sentry) applied to serverless, so an agent receives a scoped root-cause unit instead of having to reconstruct the failure from scattered logs. (Source: sources/2026-09-30-cloudflare-detect-and-send-production-issues-straight-to-your-agent)
- Zero-instrumentation because detection lives in the runtime. Issues is
"built into the Workers runtime, so there is no SDK to install or application
wrapper to add." One line —
observability.issues.enabled = trueinwrangler.jsonc— turns it on. It records uncaught exceptions, failed invocations, HTTP 5xx,console.log()/console.error()output, and logs containing a stack trace; it also flags runaway alarm conditions and code that writes large log volumes inside loops. Contrast with the classic install-an-SDK error-tracking model. (Source: sources/2026-09-30-cloudflare-detect-and-send-production-issues-straight-to-your-agent) - The platform captures what happened; the app must add who it happened to.
Cloudflare can see the exception but not which users/accounts/sessions matter.
Teams enrich occurrences with the Workers runtime's built-in
OpenTelemetry API —
tracing.getActiveSpan()?. setAttribute("user.id", …)— with no extra package. Those identifiers then appear on every occurrence, so you can see whether failures concentrate in one account or session before sending to an agent. (Source: sources/2026-09-30-cloudflare-detect-and-send-production-issues-straight-to-your-agent) - The handoff is a threshold-triggered automation, not a dashboard task. "An issue no longer has to sit in a dashboard while someone copies a stack trace and pastes it into a prompt." Configure an Automation once; when an issue crosses an occurrence threshold or returns after a quiet period, Issues sends the failure summary + diagnostic context straight to the destination. This is the detection = dispatcher half of the Dispatcher–Coding-Agent–Closer loop — the trigger is a live error signal rather than a ticket query. (Source: sources/2026-09-30-cloudflare-detect-and-send-production-issues-straight-to-your-agent)
- Three handoff destinations. (a) Built-in coding agents — Claude Code (routine ID + token), Cursor (automation webhook URL), Devin (API token + org ID); (b) generic webhooks — POST issue context to your own agent or HTTPS endpoint; (c) chat / incident management — notify a team or on-call workflow. The payload: exception, error, source-mapped stack trace, leading and trailing logs/traces, Worker version, and the application context you added. (Source: sources/2026-09-30-cloudflare-detect-and-send-production-issues-straight-to-your-agent)
- Deeper investigation is delegated to MCP, not stuffed into the payload.
The automation ships a bounded diagnostic snapshot; for open-ended digging the
agent is connected separately to
Cloudflare's MCP server
(
workers-observability) so it can query the related logs and traces on demand, then propose code + test changes and open a PR. Separating the fixed handoff payload from the pull-based MCP query surface keeps the trigger cheap and the investigation unbounded. (Source: sources/2026-09-30-cloudflare-detect-and-send-production-issues-straight-to-your-agent) - The human stays at the merge gate. "You stay in control of what reaches production: review the pull request, deploy the fix, and mark the issue resolved." The agent prepares the change; a human authorizes it — the human-in-the-loop merge gate that distinguishes this from fully autonomous closed-loop remediation. (Source: sources/2026-09-30-cloudflare-detect-and-send-production-issues-straight-to-your-agent)
- Dogfooded on Workflows — two bugs in one day. Cloudflare turned Issues on for Workflows (itself built entirely on the Workers platform) and within a day found two edge-case bugs hidden in a large volume of traffic: (a) a control-plane migration stuck in a retry loop on a SQLite foreign-key error, and (b) a deletion process that never completed because deleting Workflow instances could exceed a Workers subrequest limit. The automation routed both issues into Cloudflare OS, which followed the errors into the Workflows code and proposed a fix for each — "instead of leaving the team to connect thousands of separate pieces of telemetry and user reports." (Source: sources/2026-09-30-cloudflare-detect-and-send-production-issues-straight-to-your-agent)
Systems / concepts / patterns extracted¶
- Systems: Cloudflare Workers (runtime the
detection is built into); Cloudflare
Workflows (dogfood target — two production bugs found);
Cloudflare OS (agent destination that followed the
errors into code and proposed fixes); OpenTelemetry
(runtime built-in tracing API used to attach app context —
user.id,account.id,session.id— to occurrences); MCP (Cloudflare'sworkers-observabilityMCP server for on-demand log/trace querying by the agent); cf CLI (an agent can use it to inspect issues and create automations). - "Issues" (the product itself): recorded here as prose + tags, not as a standalone wiki page — it is a feature of Workers observability, single-source today. Leave for Lint promotion if it earns ≥2 distinct sources.
- Concepts: Observability (error monitoring as a zero-instrumentation, grouped, agent-consumable layer); Human-in-the-loop (review the PR before deploy); Alert fatigue (grouping repeated exceptions into one issue is the deduplication answer).
- Patterns: Dispatcher–Coding-Agent–Closer (Issues = the detection-triggered dispatcher; the coding agent opens a PR; the human closes) and closed-loop remediation (the human-out-of-the-loop sibling, held apart here because the fix path keeps a human merge gate).
Operational numbers / specifics¶
- Config:
observability.issues.enabled = trueinwrangler.jsonc(one line, no SDK, no wrapper). - What Issues records: uncaught exceptions, failed invocations, HTTP 5xx,
console.log()/console.error()output, logs containing a stack trace; flags runaway alarm conditions and large log volumes written inside loops. - Grouping surface per issue: the error, stack trace (when available), leading/trailing logs and traces, Worker version, request details, first-seen, occurrence count, and trend over time.
- Automation triggers: occurrence-threshold crossing, or return-after-a-quiet-period.
- Handoff payload: exception, error, source-mapped stack trace, leading and trailing logs and traces, Worker version, user-added application context.
- Built-in agent integrations: Claude Code (routine ID + token), Cursor (automation webhook URL), Devin (API token + organization ID); plus generic webhooks and chat/incident tools.
- MCP:
github.com/cloudflare/mcp-server-cloudflare→apps/workers-observability. - Dogfood result: two Workflows bugs found and fixed within one day of enabling (SQLite FK migration retry loop; subrequest-limit-exceeded deletion).
Caveats¶
- Launch/product-announcement post. This is an "introducing Issues (open beta)" post; it clears the AGENTS.md borderline bar because the observability architecture (runtime-built-in detection, error grouping, threshold-triggered structured handoff) and the two dogfooded production incidents are the substantive core — but there are no hard SLO/latency/scale numbers for Issues itself (no MTTD/MTTR figures, no detection-latency, no throughput).
- The two Workflows incidents are described qualitatively ("within a day," "an edge case") with no timeline or magnitude beyond root cause.
- Agent-proposed fixes are explicitly gated on human PR review + deploy; the post does not claim autonomous merge.
Source¶
- Original: https://blog.cloudflare.com/real-time-issue-detection/
- Raw markdown:
raw/cloudflare/2026-09-30-detect-and-send-production-issues-straight-to-your-agent-7fbfd064.md
Related¶
- systems/cloudflare-workers — Issues is built into this runtime.
- systems/cloudflare-workflows — dogfood target; two production bugs found.
- systems/cloudflare-os — agent destination that proposed the fixes.
- systems/opentelemetry — runtime built-in API for app-context enrichment.
- systems/model-context-protocol — the agent's pull-based log/trace query surface.
- systems/cf-cli — inspect issues + create automations from the CLI.
- concepts/observability — the discipline this implements at the error layer.
- concepts/human-in-the-loop — the PR-review merge gate.
- patterns/dispatcher-coding-agent-closer — Issues = the detection-triggered dispatcher.
- patterns/closed-loop-remediation — the human-out-of-the-loop sibling.
- companies/cloudflare