Identify AI model overuse with User Insights¶
Summary¶
Cloudflare extends AI Gateway User Insights — the identity-aware analytics tab that shows which users, apps, tasks, and models drive AI traffic — with task and model-fit context. The headline addition is a model overkill view that flags conversations where the selected model is more capable (and costlier) than the task requires (e.g. a summarization request sent to a large reasoning model). Supporting this are task analysis (grouping conversations into coding / research / writing / summarization / data-analysis categories), turns analysis (how much back-and-forth a task takes, and the cumulative time/token/cost across turns), and a new Potential Savings view. The same task and conversation signals power the Auto Router (launching in beta alongside this release). Architecturally, classification is done by a dedicated Cloudflare Worker that runs asynchronously, off the request path over already-proxied AI Gateway logs — it examines the full conversation trajectory (user requests, assistant responses, tool calls, tool results), emits a task category plus a confidence score and per-conversation dimensions (complexity, intent ambiguity, stakes, context dependence), and joins the result back to log metadata. It follows AI Gateway's existing log architecture (metadata in Durable Objects, log bodies in R2), which is what keeps classification out of the hot path — at the cost of the dashboard trailing incoming traffic by ~1 day.
Key takeaways¶
- Token counts and model names don't tell you what work is being done. The same token count can be a code review, a research task, or an agent making several calls — you can't evaluate model choice without understanding the task. User Insights now adds that missing context. (Source: this article)
- Model overkill view = model-fit signal, not a leaderboard. It flags conversations where a high-capability model was used for a simple task, attributes the pattern to specific users/agents/applications, and lets teams compare cost, latency, token usage, and conversation turns for the same task type. It explicitly does not auto-recommend a replacement model — it surfaces better questions ("would a faster/cheaper model produce an equivalent outcome?").
- Common root causes of overkill are organizational, not technical: a model is the default, users are unsure which model to pick, or an agent is configured to use the same model for every step.
- Task analysis groups conversations into initial categories: coding, research, writing, summarization, data analysis. This reveals that "a surprising amount of traffic comes from simple tasks even though those tasks are being sent to a high-capability model."
- Turns analysis exposes the full cost of a task: the first request is only part of it. A simple task that keeps taking several turns is a prompt / model / workflow smell. A long conversation on complex work is not inherently bad.
- Classification is asynchronous and off the request path. A dedicated Worker processes eligible logs after AI Gateway has served the response, so it adds no latency to the user's response. The tradeoff: User Insights is not real-time — analysis may trail incoming traffic by ~1 day. It's a pattern-over-time tool, not a live monitor. (Source: this article)
- The classifier reuses AI Gateway's split-storage log architecture: metadata stored separately from log bodies, with the current implementation using Durable Objects for metadata and R2 for log bodies. User Insights exposes derived categories and aggregate views, not a raw prompt browser; log-body retention still follows the configured AI Gateway logging behavior.
- These signals feed the Auto Router (closed beta, same release). The Auto Router uses conversation trajectory, task category, task complexity, and model-fit signals to route each request to an appropriate model taking cost into account — so a team doesn't need a separate routing rule per workload. It does not always pick the cheapest model.
- Identity-awareness comes from AI Gateway sitting behind
Cloudflare Access. Authenticated users and
sessions map to AI traffic — for teams' own apps and for developer
tools/agent harnesses like Claude Code, Codex, and OpenCode, which inherit
the identity context automatically. Custom apps must pass a stable
user_idandsession_id(viacf-aig-metadata) — non-sensitive identifiers kept out of the prompt itself. Access is free for teams up to 50 users. - Free to AI Gateway users. The capabilities are available at no cost.
Extracted architecture¶
Classification pipeline (the substantive infra content):
- Categorization engine = a dedicated Cloudflare Worker (not inline middleware) that processes eligible AI Gateway logs.
- Input: the whole conversation trajectory — user requests, assistant responses, tool calls, tool results.
- Output: a task category (coding, debugging, research, summarization, …) + a confidence score + evaluated dimensions: task complexity, intent ambiguity, stakes, context dependence. The category is joined with the log metadata the dashboard uses.
- Placement: asynchronous, out of the request path. AI Gateway writes the log to its existing storage path first; the classifier processes it afterward. This is a deliberate latency-tolerant design — zero added serving latency, at the cost of a ~1-day dashboard lag.
- Storage: follows AI Gateway's existing log architecture — metadata separate from log bodies, currently Durable Objects (metadata) + R2 (log bodies) (a log-body vs derived-view split; the derived categories/aggregates are what the dashboard renders, not raw prompts).
- Scope discipline: the signal is "used for reporting and routing analysis, and is not intended to replace or expose the original request." Small, understandable category set by design, rather than inferring every detail of a user's work.
Views layered on top: model overkill view, task analysis (category breakdown), turns analysis (per-session turn count + cumulative time/tokens/cost), Potential Savings view.
Identity plumbing: Access in front of AI
Gateway attaches verified user/session identity; agent harnesses (Claude Code,
Codex, OpenCode) inherit it; custom apps pass cf-aig-metadata with user_id,
session_id, idp_group, application.
Operational numbers¶
- Analysis lag: ~1 day (classification trails incoming traffic as logs are processed and aggregated).
- Cloudflare Access: free for teams up to 50 users.
- Task categories (initial): 5 — coding, research, writing, summarization, data analysis.
- Per-conversation scored dimensions: 4 — complexity, intent ambiguity, stakes, context dependence (plus a confidence score on the category).
- Pricing: free to AI Gateway users.
Caveats¶
- This is a product-update post; it is light on the internals of the classifier model itself (which model powers the categorization Worker, its accuracy/precision, sampling policy, what makes a log "eligible", exact session-boundary rules, retention of the derived signals).
- "~1 day" is stated as approximate, not a documented SLA.
- The storage claim is scoped: "the current implementation uses Durable Objects for metadata and R2 for log bodies" — explicitly a current-implementation detail, not a stable contract.
- Model overkill is a starting point for investigation, not an automated action; the post is careful that it neither recommends replacements nor blocks traffic. (The automated-action counterpart is the separately-shipping Auto Router.)
Source¶
- Original: https://blog.cloudflare.com/ai-model-overuse-user-insights/
- Raw markdown:
raw/cloudflare/2026-09-30-identify-ai-model-overuse-with-user-insights-9063144a.md
Related¶
- systems/cloudflare-ai-gateway-user-insights — the feature this article extends.
- systems/cloudflare-ai-gateway — the proxy tier User Insights observes.
- concepts/model-first-routing — the routing model these task signals feed (via Auto Router).
- concepts/efficiency-frontier — model overkill = picking above the efficiency frontier.
- patterns/central-proxy-choke-point — why the proxy can classify all traffic out-of-band.
- patterns/cheap-approximator-with-expensive-fallback — the Auto Router's cost/capability tiering economics.
- companies/cloudflare