Skip to content

CONCEPT Cited by 4 sources

Efficiency frontier

Definition

The efficiency frontier is the set of models that offer the best price point for a given level of intelligence. It is deliberately contrasted with the intelligence frontier — the set of highest-intelligence models that frontier labs race to advance (which can now solve novel math or cybersecurity problems).

The key claim: when AI is deployed at scale, the efficiency frontier matters more than the intelligence frontier, because most day-to-day coding does not require peak intelligence — it requires a model that clears the quality bar for typical software-engineering work at the lowest cost. (Source: sources/2026-08-07-databricks-managing-ai-coding-costs-at-scale)

Why it matters

  • It advances faster than the intelligence frontier. New models present better intelligence-per-unit-price "almost weekly." So the largest single cost lever available to an organization is rapidly moving spend onto newer, more efficient models as they ship — not chasing peak intelligence.
  • "Cheaper model" is not a simple axis. Cost and quality are entangled: a model can be cheaper per token yet more expensive per completed task (more retries, more tokens, worse cache behavior). The efficiency frontier is about cost at a fixed quality bar, measured on realistic tasks — not sticker price per token.

Capturing the frontier requires knowing your incumbents

Public benchmarks poorly predict real-world coding performance, so organizations build internal automated evaluations representative of their own development mix to decide which new models actually beat incumbents. Two consequences:

  • Positive result → roll out. Databricks benchmarked models on its multi-million-line codebase, found GLM models highly competitive on price/performance, and rolled GLM out to developers internally.
  • Negative results are common. Stripe found Opus 4.7 did not meaningfully improve quality over Opus 4.6 while costing more, and declined to make it available; Databricks saw a cost regression comparing Opus 5.0 to 4.8. Many new models do not advance the efficiency frontier.

Architectural implications

Chasing the efficiency frontier only pays off if the surrounding infrastructure lets you switch models cheaply:

Seen in

  • sources/2026-09-28-databricks-how-databricks-rolls-out-frontier-models-to-12000-employees — the frontier as a promote/drop gate on Day-1 model rollout. Databricks operationalises "is this model actually on the efficiency frontier?" as the decision that promotes a new model into standard circulation or drops it. Concrete negatives and positives: Opus 5.0 regressed (more expensive and lower quality than Opus 4.8) — a marketed frontier model off the actual frontier; but after Day-1 experimental exposure, Opus 5.5 came in −29% $/session vs Opus 4.8 and GPT-6 Sol −48% vs GPT-5.6 Sol (stratified re-weighted comparison) — both promoted to GA. GPT-6 Sol was not made the Codex default (occasional quality downgrade) but added to the smart router for its cost profile — a case where a model advances the cost axis of the frontier without advancing quality. Sharpens this page's "cheaper per token ≠ cheaper per task" rule with a "measure on your own stratified session distribution" method. (Source: sources/2026-09-28-databricks-how-databricks-rolls-out-frontier-models-to-12000-employees)

  • sources/2026-08-07-databricks-managing-ai-coding-costs-at-scale — coins the efficiency-frontier framing as the single greatest cost lever; cites GLM rollout and Opus negative-eval examples.

  • sources/2026-09-30-cloudflare-cut-your-ai-spend-with-ai-gateways-auto-router — the "jagged frontier" framing operationalized. Cloudflare's Auto Router (cloudflare/auto) argues "the ability to solve a problem often exists somewhere in this portfolio of models; the router's job is to choose the right model for each task while balancing quality and price," and that savings grow with how much non-frontier work you have. It also sharpens this page's "cheaper per token ≠ cheaper per task" point into a design rule: a router should minimize predicted trajectory cost, not load-balance by dollars per million tokens (a model cheaper on paper can burn disproportionately more tokens). Reported 86.6% success at ~35% of Opus cost on Cloudflare's internal benchmark. (Source: sources/2026-09-30-cloudflare-cut-your-ai-spend-with-ai-gateways-auto-router)
  • sources/2026-09-30-cloudflare-identify-ai-model-overuse-with-user-insights — "model overkill" = the visible symptom of running above the efficiency frontier. Cloudflare's User Insights model overkill view flags conversations where a high-capability model was used for a simple task and attributes them to specific users/agents/apps — making "you're paying for intelligence the task didn't need" an inspectable, per-identity signal rather than an abstract framing. It reframes the frontier as an operational monitoring problem: task analysis + turns analysis let a team compare cost/latency/tokens/turns for the same task type and ask "would a cheaper capable model produce an equivalent outcome?" — the efficiency- frontier question, per workload. (Source: sources/2026-09-30-cloudflare-identify-ai-model-overuse-with-user-insights)

Merged aliases

  • capacity-efficiency
  • price-performance-ratio
Last updated · 766 distilled / 2,225 read