Skip to content

SYSTEM Cited by 3 sources

Cloudflare AI Search

Overview

AI Search (formerly AutoRAG) is Cloudflare's managed search primitive for AI agents: hybrid BM25 + vector retrieval on built-in storage + vector index, where each search instance is dynamically creatable and destroyable at runtime from a Worker via the new ai_search_namespaces binding.

Cloudflare's positioning in the 2026-04-16 launch post:

"If you're building search yourself, you need a vector index, an indexing pipeline that parses and chunks your documents, and something to keep the index up to date when your data changes. If you also need keyword search, that's a separate index and fusion logic on top. And if each of your agents needs its own searchable context, you're setting all of that up per agent. AI Search is the plug-and-play search primitive you need."

— (Cloudflare, 2026-04-16)

Architectural placement

AI Search sits across four of Cloudflare's existing primitives:

  • R2 — managed storage substrate per instance; optional external R2 bucket as a data source.
  • Vectorize — vector index substrate per instance.
  • Browser Run (formerly Browser Rendering) — built-in website crawler when a website is the data source, now bundled into AI Search rather than billed separately.
  • Workers + Durable Objects — the consumer; ai_search_namespaces binding is invoked from Worker code, typically within a DO-hosted agent like the 2026-04-16 post's SupportAgent worked example.

The instance — (storage, vector-index, BM25-index, indexing-pipeline, query-pipeline, optional-external-source) — is the unit. A namespace groups instances and is the scope at which runtime creation, deletion, listing, and cross-instance search happen.

Core capabilities

1. Hybrid search (new in 2026-04-16 release)

BM25 + vector running in parallel with engine-side fusion. Pre-release AI Search / AutoRAG was vector-only; 2026-04-16 adds BM25 as a first-class engine and exposes the pipeline as configurable instance options.

const instance = await env.AI_SEARCH.create({
  id: "my-instance",
  index_method: { keyword: true, vector: true },
  indexing_options: { keyword_tokenizer: "porter" },  // or "trigram" for code
  retrieval_options: { keyword_match_mode: "or" },    // or "and"
  fusion_method: "rrf",                                // or "max"
  reranking: true,
  reranking_model: "@cf/baai/bge-reranker-base"
});

All options have sane defaults. See concepts/retrieval-ranking-funnel for the primitive, concepts/retrieval-ranking-funnel for RRF, concepts/retrieval-ranking-funnel for the rerank stage.

2. Built-in storage and index

  • instance.items.uploadAndPoll(filename, content, { metadata: { … } }) — upload + index in one awaitable call; returns item.status = "completed" once searchable. See upload-then-poll-indexing.
  • No R2 bucket to pre-provision; no external data source required. One external source (R2 bucket or website) can optionally be attached alongside the built-in storage, with a sync schedule.

3. Namespace binding — runtime-provisioned instances

// wrangler.jsonc
{
  "ai_search_namespaces": [
    { "binding": "AI_SEARCH", "namespace": "example" }
  ]
}

Surface: env.AI_SEARCH.create(…), env.AI_SEARCH.delete(…), env.AI_SEARCH.list(…), env.AI_SEARCH.search(…), env.AI_SEARCH.get(id) → instance handle for per-instance items.uploadAndPoll() / search().

Replaces the previous env.AI.autorag() API that accessed AI Search via the AI binding; old bindings continue to work through Workers compatibility dates.

Canonical wiki instance of runtime-provisioned-per-tenant-search-index and the retrieval-tier realisation of one-to-one-agent-instance.

4. Metadata boost at query time

const results = await instance.search({
  query: "deployment guide",
  ai_search_options: {
    boost_by: [{ field: "timestamp", direction: "desc" }]
  }
});

timestamp is built into every item; any custom metadata field (priority, region, language, tenant, …) defined at indexing time can also drive a boost. Business logic layered on top of relevance, not fused into it. See metadata-boost, metadata-boost-at-query-time.

const results = await env.SUPPORT_KB.search({
  query: "billing error",
  ai_search_options: {
    instance_ids: ["product-knowledge", "customer-abc123"]
  }
});

Merges + ranks across instances in a single call. Namespace-level generalisation of patterns/tool-surface-minimization; see patterns/parallel-retrieval-fusion.

Canonical usage shape — support agent

The 2026-04-16 post walks through a customer-support agent built on the Agents SDK:

namespace: "support"
├── product-knowledge     (R2 as source, shared across all agents)
├── customer-abc123       (managed storage, per-customer)
├── customer-def456       (managed storage, per-customer)
└── customer-ghi789       (managed storage, per-customer)
  • One shared product-knowledge instance, R2-backed, contains product docs across all agents.
  • One per-customer instance, managed-storage, accumulates past-resolution summaries as agent memory.
  • SupportAgent.onChatMessage creates the per-customer instance on first appearance (idempotent — try { … } catch {}).
  • Two tools exposed to the model:
  • search_knowledge_base — fans across product-knowledge + customer-<id> in one call, with boost_by: timestamp to surface recent docs.
  • save_resolution — instance.items.uploadAndPoll(filename, content) on resolution, so future agents see it.
  • LLM: Kimi K2.5 via Workers AI (@cf/moonshotai/kimi-k2.5).
  • Durable-object backing for conversation state via AIChatAgent from the Agents SDK.
  • stepCountIs(10) caps agentic tool-use loops.

DX improvements + Dev Stack MCP (2026-08-06)

The 2026-08-06 "give your agents a search engine for your data" post ships developer-experience refinements that eliminate the manual stitching of Workers AI + AI Gateway + Vectorize + R2 + Browser Run — "now, AI Search can do this automatically — and better."

Sitemap-less crawling via --parse-type discover

Previously the website integration required a sitemap. The Discover parsing option adds a website without one by following links, powered by Browser Run's /crawl:

npx wrangler ai-search instance create cloudflare-community \
  --namespace dev-stack \
  --source https://community.cloudflare.com \
  --type web-crawler \
  --parse-type discover

Public /search + /mcp endpoints (no code)

Enabling public URLs on a namespace immediately yields /search and /mcp endpoints that query every instance in the namespace, with no auth and nothing to deploy — the zero-code alternative to a Worker binding. Canonical instance of public-endpoint-no-code-deployment.

"Reach for the Worker when you're folding search into an existing app or MCP server, as we are. Or reach for the public endpoint when you just want a shareable search endpoint in one click."

— (Cloudflare, 2026-08-06)

Custom domains + Cloudflare Access

Public endpoints can be branded with a custom domain (search.example.com/mcp) and locked down with Cloudflare Access to make a private search instance requiring login for any human or agent. See custom-domain-over-public-endpoint.

Cloudflare Dev Stack MCP — the canonical multi-instance instance

The Cloudflare Dev Stack MCP (playground.ai.cloudflare.com) gives coding agents "current, cited docs from across the Cloudflare developer ecosystem, so they build on the latest features and fixes instead of stale training data." It is built as follows:

  1. Index each surface — one AI Search instance per Cloudflare-owned surface: Docs, Blog, API Docs, Community, Astro, Vite, Vitest, Hono, Replicate, OpenNext — all in a single dev-stack namespace. "Because Cloudflare owns the website data, AI Search is able to treat them as a single set and ingest them all the same way."
  2. Combine into one search — a single Worker-hosted MCP tool makes one multi-instance call across all 10 instances (instance_ids: [...], reranking: { enabled: true }), returning cited chunks tagged with the instance they came from (patterns/parallel-retrieval-fusion). Ships as a tool inside Cloudflare's MCP server.
  3. Brand + lock down — custom domain + optional Cloudflare Access.

wrangler.jsonc binding for the namespace:

{
  "ai_search_namespaces": [
    { "binding": "AI_SEARCH", "namespace": "cloudflare-stack" }
  ]
}

Drop into any coding agent's MCP config:

{
  "mcpServers": {
    "dev-stack": { "url": "https://stack.mcp.cloudflare.com/mcp" }
  }
}

"That replaces the usual fallback (web search then fetching full pages), which is slow, token-heavy, and often lands on the wrong or stale source."

Bot policy compliance

AI Search's crawler identifies as Cloudflare-AI-Search with an immutable public user agent, follows robots.txt, and respects whatever bot controls a site has in place — same posture as Browser Run.

Preview pricing (2026-08-06)

Free during beta; billing not yet enabled. As-modelled preview:

Category Price Free monthly allotment
Base ingestion $0.75 / 1M tokens 5M tokens †
Image processing (add-on) +$0.50 / 1M tokens (shares 5M pool)
Stored data $2.00 / GB-month 10 GB
Semantic query (hybrid + vector) $0.75 / 1K queries 2,000 queries ‡
Full-text query $0.10 / 1K queries 2,000 queries ‡
Embedding + reranking Free with select Workers AI models N/A

† single 5M-token ingestion pool across all file types. ‡ single 2,000-query pool shared across both query types.

Embedding and reranking are free with default Workers AI models — "the models behind indexing and every search are not a cost you have to worry about." Answer generation and query rewriting are optional, billed as Workers AI usage (or AI Gateway credits). Example bill: 20K docs (~20M text tokens) + 1K images + 30K semantic queries/month ≈ ~$35 first month (mostly one-time ingestion), dropping to ~$21/month afterward (queries-dominated). Ingestion chunking uses ~10% overlap (× 1.1 in the bill).

Cloudflare.com, Developer Docs, and Blog all run their search on AI Search (all using hybrid search — semantic + keyword in one query). Sites on EmDash can add semantic search via the AI Search plugin.

General availability — multimodal, OCR, bigger files, pricing (2026-10-01)

The 2026-10-01 "AI Search is now generally available" post takes AI Search GA and ships four substantive changes on top of the hybrid-search + DX releases: native multimodal retrieval, OCR, larger files, and finalized usage-based pricing. "The search on our own blog and developer docs" runs on it.

Native multimodal embedding and retrieval

Pre-GA image support was "naive: we would perform object detection, generate a caption, and then embed that text" — images were searchable only through what the caption captured. GA now does both: caption-based textual understanding and native image retrieval.

  • Native image embeddings embed the image pixels directly into the same vector space as indexed text, preserving visual detail (texture, document layout, chart relationships, palette, composition, spatial relationships) that a caption compresses away.
  • The native-multimodal model is Qwen3-VL-Embedding (on Workers AI). At query time AI Search checks whether the instance's embedding model supports images; if so, a query image is embedded directly by that model, landing in the same space as indexed images and text.
  • Graceful fallback for text-only embedders: if the configured model is text-only, a query image is converted to text with ToMarkdown (Browser Run) and searched by the resulting caption. "This gives every model basic multimodal support, while models with native image support get the full visual signal."
  • Enables queries hard to express in words: image→image ("find visually similar"), text→image, or combined ("a bird with similar markings"). Useful for product discovery, screenshot matching, charts/diagrams, scanned documents.

This is the retrieval realisation of multimodal embeddings in one space — a sibling of Figma AI Search's CLIP usage, but productised with a built-in text-only fallback.

Matryoshka Representation Learning keeps it cheap

Richer image embeddings are larger than caption text, so to "keep these richer representations efficient, AI Search leverages Matryoshka Representation Learning (MRL), allowing smaller embeddings to retain useful information while keeping storage manageable and search fast." MR-trained embeddings can be truncated to a prefix without re-embedding, turning dimensionality into a serving-time storage/speed knob over the Vectorize index.

OCR for scanned documents + bigger files

  • Text files (Markdown, HTML, CSV, JSON, …) and PDFs now go up to 10 MiB (from 4 MiB).
  • Many PDFs are really scanned images with no extractable text; turning on OCR makes AI Search "read the text from each page before chunking and embedding it." OCR is available to every account and billed as image-processing ingestion tokens under the new pricing.

The life of an AI Search query (query lifecycle, stated explicitly)

GA documents the end-to-end query path, which is the canonical retrieve → fuse → rerank / parallel-retrieval-fusion shape:

query
  → (optional) rewrite
  → embed  (image queries: embedded directly by multimodal model,
            or captioned first by a text-only model)
  → vector search ∥ keyword search   (run in parallel)
  → fuse  (RRF / max)
  → (optional) rerank                (cross-encoder)
  → return top chunks  OR  pass to a generation model for an answer

GA pricing (billing live 2026-11-01)

AI Search pricing is designed so you can estimate your bill before indexing a single file. You pay for three things — content ingested, data stored, queries run — and "the work in between (parsing, chunking, embedding with Workers AI models, keyword indexing, and reranking) is included." No instance hours, capacity units, or monthly minimums. Ingestion is one rate per token regardless of embedding model; switching text-only → multimodal doesn't change ingestion cost unless you also process images (add-on fee). Billing goes live 2026-11-01 (reminder email first).

Category Price Free monthly allotment
Base ingestion $0.75 / 1M tokens 5M tokens †
Image processing (add-on) +$0.50 / 1M tokens (shares 5M pool)
Stored data $2.00 / GB-month 10 GB
Semantic query (hybrid + vector) $0.75 / 1K queries 1,000 queries
Full-text query $0.10 / 1K queries 1,000 queries
Embedding + reranking Free with select Workers AI models (third-party billed separately) N/A

† single 5M-token ingestion pool across all file types (text, images). Change vs the August preview pricing: the free tier now grants 1,000 semantic + 1,000 full-text queries as two separate pools (not a shared pool of 2,000). Embedding + reranking remain free with default Workers AI models; answer generation + query rewriting are optional Workers AI usage.

Roadmap

  • Video + audio ingestion pipeline for rich-media search.
  • Keyword-engine refactor that "scales better… particularly for when you have big data stores where the current implementation has limits."
  • Simpler ways to enable AI Search / create indexes for sites already on Cloudflare, so agents can discover/consume content efficiently.

CLI surface

npx wrangler ai-search create my-search creates an instance; consistent with the cf CLI + unified TypeScript schema rollout (sources/2026-04-13-cloudflare-building-a-cli-for-all-of-cloudflare) — AI Search's ~N operations appear in the same ~3,000-operation surface exposed across CLI / bindings / MCP Code Mode / Terraform / wrangler.jsonc.

Dogfood

"The search on our blog is now powered by AI Search. Try the magnifying glass icon to the top right."

— (Cloudflare, 2026-04-16)

Third instance of the "dogfood the platform as a customer-facing product" recurring shape in April 2026, after Agent Lee and Project Think. See companies/cloudflare.

Open-beta limits (2026-04-16)

Limit Workers Free Workers Paid
AI Search instances per account 100 5,000
Files per instance 100,000 1M (500K for hybrid search)
Max file size 4 MB 4 MB
Queries per month 20,000 Unlimited
Max pages crawled per day 500 Unlimited

Note (GA, 2026-10-01): the max file size was raised to 10 MiB for text files + PDFs. See the GA section above.

Pricing: free during open beta; Browser Run website crawling bundled in (not separately billed); Workers AI + AI Gateway still billed separately. Goal post-beta: "unified pricing for AI Search as a single service."

Pre-release instances

Instances created before 2026-04-16 continue to work — customer-visible R2 buckets, Vectorize indexes, Browser Run usage remain billed as before. Migration path promised.

Caveats

  • Preview / open-beta release; SLA, cross-region behaviour, durability guarantees not disclosed.
  • No published latency, throughput, recall, or cost-per-query numbers.
  • Embedding model powering the vector half is not named.
  • Chunking strategy, chunk overlap, structured-format handling inside the indexing pipeline are opaque — the value prop is "upload and trust."
  • Cross-encoder reranker listed with one option (@cf/baai/bge-reranker-base); pluggability unclear.
  • No sparse / learned-sparse retrieval (SPLADE / ELSER style); BM25 + dense only.
  • No explicit competitive positioning vs Pinecone / Weaviate / Qdrant / Atlas Hybrid Search / pgvector+FTS.

Seen in

Last updated · 766 distilled / 2,225 read