Skip to content

SYSTEM Cited by 2 sources

Cloudflare AI Gateway User Insights

Cloudflare AI Gateway User Insights is an AI Gateway analytics feature that turns already-proxied AI traffic into a behavioral view for each identified account. It learns a per-account session-cost baseline, ranks sessions that depart from that history, and presents a review queue to an administrator. Since the 2026-09-30 update it also adds task and model-fit context: a model overkill view, task-category analysis, turns analysis, and a Potential Savings view — the same signals that power the AI Gateway Auto Router. The feature is monitoring-oriented: it does not infer intent or automatically block an account. (Source: sources/2026-08-05-cloudflare-catching-rogue-ai-behavior-with-identity-aware-analytics, sources/2026-09-30-cloudflare-identify-ai-model-overuse-with-user-insights)

Detection model

For each account, User Insights evaluates sessions rather than isolated requests:

  1. Build a rolling personal baseline from the p95 session cost over the last 30 days.
  2. Treat a session above 2× that p95 as a relative anomaly candidate.
  3. Require the candidate to also exceed the organisation's account-level p99 session-cost floor.
  4. Apply a dollar floor so a very large percentage increase on a tiny spend value does not alert.
  5. Put qualifying accounts in an administrator-facing rogue-behavior feed.

The two thresholds deliberately separate unusual behavior from material spend. A high-cost user whose sessions are normal for them should not alert merely because they are expensive; a low-spend agent whose cost grows by 10× should not disappear inside an organisation-wide dollar threshold.

Identity dependency

The strongest per-user view depends on Cloudflare Access protecting a custom AI Gateway domain. Access authenticates the caller through a SAML-supported identity provider and AI Gateway attaches the verified Access user ID as cf.user_id request metadata. Without that mapping, User Insights can still observe account IDs, but cannot turn the anomalous entity into a person or agent an administrator can readily investigate.

Scope and controls

User Insights tracks cost signals and highlights associated inefficiency hints such as low cache-hit rate or oversized context windows. It does not itself decide whether use is malicious, whether a credential is compromised, or whether a request should be stopped. Separate AI Gateway spend controls can block additional requests or route to a lower-cost model after a per-user limit is reached.

Model fit and task context (2026-09-30 update)

The 2026-09-30 update widens User Insights from a spend-anomaly view to a model-fit view. It adds four related surfaces (Source: sources/2026-09-30-cloudflare-identify-ai-model-overuse-with-user-insights):

  • Model overkill view. Flags conversations where the selected model appears more capable — and costlier — than the task requires (e.g. a formatting / summarization request sent to a high-capability reasoning model). It attributes the pattern to specific users, agents, or applications and lets a team compare latency, input/output tokens, conversation turns, and total cost for the same task type. Deliberately not a leaderboard and it does not auto-recommend a replacement model — it surfaces investigative questions ("would a faster/cheaper model produce an equivalent outcome? is the extra capability improving the result?"). This is the picking-above-the- efficiency frontier signal, made visible per user/agent.
  • Task analysis. Groups conversations into categories — initially coding, research, writing, summarization, data analysis — so a list of model names gains the "what work is this?" context it otherwise lacks. Frequently reveals that a surprising amount of traffic is simple tasks routed to a high-capability model.
  • Turns analysis. Shows how much back-and-forth a task takes and the cumulative time / tokens / cost across turns — the first request is only part of the cost. A long conversation on complex work is fine; a simple task that keeps taking several turns is a prompt/model/workflow smell.
  • Potential Savings view. Surfaces requests that a faster or cheaper model could likely handle without compromising output quality.

Common root causes of overkill are organizational, not technical: the model is the default, users are unsure which model to pick, or an agent is configured to use the same model for every step.

Classification pipeline (how the signals are produced)

The task/model-fit signals come from a dedicated Cloudflare Worker — the categorization engine — not from inline request middleware (Source: sources/2026-09-30-cloudflare-identify-ai-model-overuse-with-user-insights):

  • Input: the whole conversation trajectory (user requests, assistant responses, tool calls, tool results).
  • Output: a task category (coding, debugging, research, summarization, …), a confidence score, and four scored dimensions — task complexity, intent ambiguity, stakes, context dependence. The category joins back to the log metadata the dashboard uses, and can also be used to evaluate model fit by comparing candidate models' suitability against their cost.
  • Placement — asynchronous, off the request path. AI Gateway writes the log to its existing storage path first; the classifier processes it afterward. This adds no latency to the user's response — a deliberately latency-tolerant design. The tradeoff: User Insights is not real-time and may trail incoming traffic by ~1 day; it's a pattern-over-time tool, not a live monitor.
  • Storage. Follows AI Gateway's existing log architecture — metadata stored separately from log bodies, currently Durable Objects for metadata and R2 for log bodies. User Insights exposes derived categories and aggregate views, not a raw prompt browser; underlying log-body retention still follows the configured AI Gateway logging behavior.
  • Feeds the Auto Router. The same conversation trajectory + task category + complexity + model-fit signals power the AI Gateway Auto Router (closed beta, same release), which turns these read-only insights into an automated per-request routing decision that takes cost into account (see model-first routing and task-difficulty model tiering). User Insights is the observe half; the Auto Router is the act half.

Caveats

The source does not disclose session-boundary rules, sampling or retention policy, minimum history needed for a baseline, cross-tenant aggregation boundaries, detector precision/recall, alert latency, or how frequently the rolling statistics are recomputed. Its $200 account-p99 example is Cloudflare internal traffic, not a documented default. The 2026-09-30 update does not disclose which model powers the categorization Worker, its classification accuracy, what makes a log "eligible", or the retention of the derived category signals; the "~1 day" analysis lag and the "Durable Objects for metadata + R2 for log bodies" storage split are both stated as current-implementation details, not stable contracts.

Seen in

Last updated · 766 distilled / 2,225 read