SYSTEM Cited by 3 sources
Cloudflare AI Search¶
Overview¶
AI Search (formerly AutoRAG) is Cloudflare's managed search primitive for AI agents: hybrid BM25 + vector retrieval on built-in storage + vector index, where each search instance is dynamically creatable and destroyable at runtime from a Worker via the new ai_search_namespaces binding.
Cloudflare's positioning in the 2026-04-16 launch post:
"If you're building search yourself, you need a vector index, an indexing pipeline that parses and chunks your documents, and something to keep the index up to date when your data changes. If you also need keyword search, that's a separate index and fusion logic on top. And if each of your agents needs its own searchable context, you're setting all of that up per agent. AI Search is the plug-and-play search primitive you need."
Architectural placement¶
AI Search sits across four of Cloudflare's existing primitives:
- R2 — managed storage substrate per instance; optional external R2 bucket as a data source.
- Vectorize — vector index substrate per instance.
- Browser Run (formerly Browser Rendering) — built-in website crawler when a website is the data source, now bundled into AI Search rather than billed separately.
- Workers + Durable Objects — the consumer;
ai_search_namespacesbinding is invoked from Worker code, typically within a DO-hosted agent like the 2026-04-16 post'sSupportAgentworked example.
The instance — (storage, vector-index, BM25-index, indexing-pipeline, query-pipeline, optional-external-source) — is the unit. A namespace groups instances and is the scope at which runtime creation, deletion, listing, and cross-instance search happen.
Core capabilities¶
1. Hybrid search (new in 2026-04-16 release)¶
BM25 + vector running in parallel with engine-side fusion. Pre-release AI Search / AutoRAG was vector-only; 2026-04-16 adds BM25 as a first-class engine and exposes the pipeline as configurable instance options.
const instance = await env.AI_SEARCH.create({
id: "my-instance",
index_method: { keyword: true, vector: true },
indexing_options: { keyword_tokenizer: "porter" }, // or "trigram" for code
retrieval_options: { keyword_match_mode: "or" }, // or "and"
fusion_method: "rrf", // or "max"
reranking: true,
reranking_model: "@cf/baai/bge-reranker-base"
});
All options have sane defaults. See concepts/retrieval-ranking-funnel for the primitive, concepts/retrieval-ranking-funnel for RRF, concepts/retrieval-ranking-funnel for the rerank stage.
2. Built-in storage and index¶
instance.items.uploadAndPoll(filename, content, { metadata: { … } })— upload + index in one awaitable call; returnsitem.status = "completed"once searchable. See upload-then-poll-indexing.- No R2 bucket to pre-provision; no external data source required. One external source (R2 bucket or website) can optionally be attached alongside the built-in storage, with a sync schedule.
3. Namespace binding — runtime-provisioned instances¶
// wrangler.jsonc
{
"ai_search_namespaces": [
{ "binding": "AI_SEARCH", "namespace": "example" }
]
}
Surface: env.AI_SEARCH.create(…), env.AI_SEARCH.delete(…), env.AI_SEARCH.list(…), env.AI_SEARCH.search(…), env.AI_SEARCH.get(id) → instance handle for per-instance items.uploadAndPoll() / search().
Replaces the previous env.AI.autorag() API that accessed AI Search via the AI binding; old bindings continue to work through Workers compatibility dates.
Canonical wiki instance of runtime-provisioned-per-tenant-search-index and the retrieval-tier realisation of one-to-one-agent-instance.
4. Metadata boost at query time¶
const results = await instance.search({
query: "deployment guide",
ai_search_options: {
boost_by: [{ field: "timestamp", direction: "desc" }]
}
});
timestamp is built into every item; any custom metadata field (priority, region, language, tenant, …) defined at indexing time can also drive a boost. Business logic layered on top of relevance, not fused into it. See metadata-boost, metadata-boost-at-query-time.
5. Cross-instance search¶
const results = await env.SUPPORT_KB.search({
query: "billing error",
ai_search_options: {
instance_ids: ["product-knowledge", "customer-abc123"]
}
});
Merges + ranks across instances in a single call. Namespace-level generalisation of patterns/tool-surface-minimization; see patterns/parallel-retrieval-fusion.
Canonical usage shape — support agent¶
The 2026-04-16 post walks through a customer-support agent built on the Agents SDK:
namespace: "support"
├── product-knowledge (R2 as source, shared across all agents)
├── customer-abc123 (managed storage, per-customer)
├── customer-def456 (managed storage, per-customer)
└── customer-ghi789 (managed storage, per-customer)
- One shared
product-knowledgeinstance, R2-backed, contains product docs across all agents. - One per-customer instance, managed-storage, accumulates past-resolution summaries as agent memory.
SupportAgent.onChatMessagecreates the per-customer instance on first appearance (idempotent —try { … } catch {}).- Two tools exposed to the model:
search_knowledge_base— fans acrossproduct-knowledge+customer-<id>in one call, withboost_by: timestampto surface recent docs.save_resolution—instance.items.uploadAndPoll(filename, content)on resolution, so future agents see it.- LLM: Kimi K2.5 via Workers AI (
@cf/moonshotai/kimi-k2.5). - Durable-object backing for conversation state via
AIChatAgentfrom the Agents SDK. stepCountIs(10)caps agentic tool-use loops.
DX improvements + Dev Stack MCP (2026-08-06)¶
The 2026-08-06 "give your agents a search engine for your data" post ships developer-experience refinements that eliminate the manual stitching of Workers AI + AI Gateway + Vectorize + R2 + Browser Run — "now, AI Search can do this automatically — and better."
Sitemap-less crawling via --parse-type discover¶
Previously the website integration required a sitemap. The Discover parsing option adds a website without one by following links, powered by Browser Run's /crawl:
npx wrangler ai-search instance create cloudflare-community \
--namespace dev-stack \
--source https://community.cloudflare.com \
--type web-crawler \
--parse-type discover
Public /search + /mcp endpoints (no code)¶
Enabling public URLs on a namespace immediately yields /search and /mcp endpoints that query every instance in the namespace, with no auth and nothing to deploy — the zero-code alternative to a Worker binding. Canonical instance of public-endpoint-no-code-deployment.
"Reach for the Worker when you're folding search into an existing app or MCP server, as we are. Or reach for the public endpoint when you just want a shareable search endpoint in one click."
Custom domains + Cloudflare Access¶
Public endpoints can be branded with a custom domain (search.example.com/mcp) and locked down with Cloudflare Access to make a private search instance requiring login for any human or agent. See custom-domain-over-public-endpoint.
Cloudflare Dev Stack MCP — the canonical multi-instance instance¶
The Cloudflare Dev Stack MCP (playground.ai.cloudflare.com) gives coding agents "current, cited docs from across the Cloudflare developer ecosystem, so they build on the latest features and fixes instead of stale training data." It is built as follows:
- Index each surface — one AI Search instance per Cloudflare-owned surface: Docs, Blog, API Docs, Community, Astro, Vite, Vitest, Hono, Replicate, OpenNext — all in a single
dev-stacknamespace. "Because Cloudflare owns the website data, AI Search is able to treat them as a single set and ingest them all the same way." - Combine into one search — a single Worker-hosted MCP tool makes one multi-instance call across all 10 instances (
instance_ids: [...],reranking: { enabled: true }), returning cited chunks tagged with the instance they came from (patterns/parallel-retrieval-fusion). Ships as a tool inside Cloudflare's MCP server. - Brand + lock down — custom domain + optional Cloudflare Access.
wrangler.jsonc binding for the namespace:
Drop into any coding agent's MCP config:
"That replaces the usual fallback (web search then fetching full pages), which is slow, token-heavy, and often lands on the wrong or stale source."
Bot policy compliance¶
AI Search's crawler identifies as Cloudflare-AI-Search with an immutable public user agent, follows robots.txt, and respects whatever bot controls a site has in place — same posture as Browser Run.
Preview pricing (2026-08-06)¶
Free during beta; billing not yet enabled. As-modelled preview:
| Category | Price | Free monthly allotment |
|---|---|---|
| Base ingestion | $0.75 / 1M tokens | 5M tokens † |
| Image processing (add-on) | +$0.50 / 1M tokens | (shares 5M pool) |
| Stored data | $2.00 / GB-month | 10 GB |
| Semantic query (hybrid + vector) | $0.75 / 1K queries | 2,000 queries ‡ |
| Full-text query | $0.10 / 1K queries | 2,000 queries ‡ |
| Embedding + reranking | Free with select Workers AI models | N/A |
† single 5M-token ingestion pool across all file types. ‡ single 2,000-query pool shared across both query types.
Embedding and reranking are free with default Workers AI models — "the models behind indexing and every search are not a cost you have to worry about." Answer generation and query rewriting are optional, billed as Workers AI usage (or AI Gateway credits). Example bill: 20K docs (~20M text tokens) + 1K images + 30K semantic queries/month ≈ ~$35 first month (mostly one-time ingestion), dropping to ~$21/month afterward (queries-dominated). Ingestion chunking uses ~10% overlap (× 1.1 in the bill).
Platform surfaces powered by AI Search¶
Cloudflare.com, Developer Docs, and Blog all run their search on AI Search (all using hybrid search — semantic + keyword in one query). Sites on EmDash can add semantic search via the AI Search plugin.
General availability — multimodal, OCR, bigger files, pricing (2026-10-01)¶
The 2026-10-01 "AI Search is now generally available" post takes AI Search GA and ships four substantive changes on top of the hybrid-search + DX releases: native multimodal retrieval, OCR, larger files, and finalized usage-based pricing. "The search on our own blog and developer docs" runs on it.
Native multimodal embedding and retrieval¶
Pre-GA image support was "naive: we would perform object detection, generate a caption, and then embed that text" — images were searchable only through what the caption captured. GA now does both: caption-based textual understanding and native image retrieval.
- Native image embeddings embed the image pixels directly into the same vector space as indexed text, preserving visual detail (texture, document layout, chart relationships, palette, composition, spatial relationships) that a caption compresses away.
- The native-multimodal model is
Qwen3-VL-Embedding(on Workers AI). At query time AI Search checks whether the instance's embedding model supports images; if so, a query image is embedded directly by that model, landing in the same space as indexed images and text. - Graceful fallback for text-only embedders: if the configured model is text-only, a query image is converted to text with ToMarkdown (Browser Run) and searched by the resulting caption. "This gives every model basic multimodal support, while models with native image support get the full visual signal."
- Enables queries hard to express in words: image→image ("find visually similar"), text→image, or combined ("a bird with similar markings"). Useful for product discovery, screenshot matching, charts/diagrams, scanned documents.
This is the retrieval realisation of multimodal embeddings in one space — a sibling of Figma AI Search's CLIP usage, but productised with a built-in text-only fallback.
Matryoshka Representation Learning keeps it cheap¶
Richer image embeddings are larger than caption text, so to "keep these richer representations efficient, AI Search leverages Matryoshka Representation Learning (MRL), allowing smaller embeddings to retain useful information while keeping storage manageable and search fast." MR-trained embeddings can be truncated to a prefix without re-embedding, turning dimensionality into a serving-time storage/speed knob over the Vectorize index.
OCR for scanned documents + bigger files¶
- Text files (Markdown, HTML, CSV, JSON, …) and PDFs now go up to 10 MiB (from 4 MiB).
- Many PDFs are really scanned images with no extractable text; turning on OCR makes AI Search "read the text from each page before chunking and embedding it." OCR is available to every account and billed as image-processing ingestion tokens under the new pricing.
The life of an AI Search query (query lifecycle, stated explicitly)¶
GA documents the end-to-end query path, which is the canonical retrieve → fuse → rerank / parallel-retrieval-fusion shape:
query
→ (optional) rewrite
→ embed (image queries: embedded directly by multimodal model,
or captioned first by a text-only model)
→ vector search ∥ keyword search (run in parallel)
→ fuse (RRF / max)
→ (optional) rerank (cross-encoder)
→ return top chunks OR pass to a generation model for an answer
GA pricing (billing live 2026-11-01)¶
AI Search pricing is designed so you can estimate your bill before indexing a single file. You pay for three things — content ingested, data stored, queries run — and "the work in between (parsing, chunking, embedding with Workers AI models, keyword indexing, and reranking) is included." No instance hours, capacity units, or monthly minimums. Ingestion is one rate per token regardless of embedding model; switching text-only → multimodal doesn't change ingestion cost unless you also process images (add-on fee). Billing goes live 2026-11-01 (reminder email first).
| Category | Price | Free monthly allotment |
|---|---|---|
| Base ingestion | $0.75 / 1M tokens | 5M tokens † |
| Image processing (add-on) | +$0.50 / 1M tokens | (shares 5M pool) |
| Stored data | $2.00 / GB-month | 10 GB |
| Semantic query (hybrid + vector) | $0.75 / 1K queries | 1,000 queries |
| Full-text query | $0.10 / 1K queries | 1,000 queries |
| Embedding + reranking | Free with select Workers AI models (third-party billed separately) | N/A |
† single 5M-token ingestion pool across all file types (text, images). Change vs the August preview pricing: the free tier now grants 1,000 semantic + 1,000 full-text queries as two separate pools (not a shared pool of 2,000). Embedding + reranking remain free with default Workers AI models; answer generation + query rewriting are optional Workers AI usage.
Roadmap¶
- Video + audio ingestion pipeline for rich-media search.
- Keyword-engine refactor that "scales better… particularly for when you have big data stores where the current implementation has limits."
- Simpler ways to enable AI Search / create indexes for sites already on Cloudflare, so agents can discover/consume content efficiently.
CLI surface¶
npx wrangler ai-search create my-search creates an instance; consistent with the cf CLI + unified TypeScript schema rollout (sources/2026-04-13-cloudflare-building-a-cli-for-all-of-cloudflare) — AI Search's ~N operations appear in the same ~3,000-operation surface exposed across CLI / bindings / MCP Code Mode / Terraform / wrangler.jsonc.
Dogfood¶
"The search on our blog is now powered by AI Search. Try the magnifying glass icon to the top right."
Third instance of the "dogfood the platform as a customer-facing product" recurring shape in April 2026, after Agent Lee and Project Think. See companies/cloudflare.
Open-beta limits (2026-04-16)¶
| Limit | Workers Free | Workers Paid |
|---|---|---|
| AI Search instances per account | 100 | 5,000 |
| Files per instance | 100,000 | 1M (500K for hybrid search) |
| Max file size | 4 MB | 4 MB |
| Queries per month | 20,000 | Unlimited |
| Max pages crawled per day | 500 | Unlimited |
Note (GA, 2026-10-01): the max file size was raised to 10 MiB for text files + PDFs. See the GA section above.
Pricing: free during open beta; Browser Run website crawling bundled in (not separately billed); Workers AI + AI Gateway still billed separately. Goal post-beta: "unified pricing for AI Search as a single service."
Pre-release instances¶
Instances created before 2026-04-16 continue to work — customer-visible R2 buckets, Vectorize indexes, Browser Run usage remain billed as before. Migration path promised.
Caveats¶
- Preview / open-beta release; SLA, cross-region behaviour, durability guarantees not disclosed.
- No published latency, throughput, recall, or cost-per-query numbers.
- Embedding model powering the vector half is not named.
- Chunking strategy, chunk overlap, structured-format handling inside the indexing pipeline are opaque — the value prop is "upload and trust."
- Cross-encoder reranker listed with one option (
@cf/baai/bge-reranker-base); pluggability unclear. - No sparse / learned-sparse retrieval (SPLADE / ELSER style); BM25 + dense only.
- No explicit competitive positioning vs Pinecone / Weaviate / Qdrant / Atlas Hybrid Search / pgvector+FTS.
Seen in¶
- sources/2026-04-16-cloudflare-ai-search-the-search-primitive-for-your-agents — launch + architecture + support-agent worked example.
- sources/2026-08-06-cloudflare-ai-search-give-your-agents-a-search-engine — DX improvements (sitemap-less
--parse-type discovercrawling, public/search+/mcpendpoints, custom domains + Cloudflare Access), the Dev Stack MCP multi-instance surface, and preview pricing. - sources/2026-10-01-cloudflare-ai-search-is-now-generally-available — GA: native multimodal image embeddings (
Qwen3-VL-Embedding) vs the old caption-only path, with a ToMarkdown fallback for text-only models; MRL to keep multimodal vectors cheap; OCR for scanned PDFs; 10 MiB files; the explicit query lifecycle; and finalized usage-based pricing (billing 2026-11-01).
Related¶
- systems/cloudflare-vectorize — vector-index substrate.
- systems/cloudflare-r2 — storage substrate + optional external data source.
- systems/cloudflare-workers — host runtime + binding layer.
- systems/cloudflare-durable-objects — typical consumer (the agent instance).
- systems/cloudflare-agents-sdk — the SDK framing the worked example.
- systems/cloudflare-browser-rendering — built-in crawler for website sources.
- systems/workers-ai — companion inference platform hosting the reranker + the application LLM + the
Qwen3-VL-Embeddingmultimodal embedder. - concepts/matryoshka-representation-learning — how GA keeps native image embeddings cheap (truncatable prefixes over Vectorize).
- concepts/vector-embedding — native image embedding vs caption-based; multimodal single-space retrieval.
- concepts/hybrid-search — the vector + keyword retrieval shape AI Search exposes.
- systems/bm25 — the lexical half of the hybrid retrieval surface.
- systems/atlas-hybrid-search — sibling productised hybrid search from the lexical-first camp.
- concepts/retrieval-ranking-funnel, concepts/retrieval-ranking-funnel, concepts/retrieval-ranking-funnel, concepts/vector-similarity-search, metadata-boost, per-tenant-search-instance, unified-storage-and-index, concepts/agent-memory, one-to-one-agent-instance.
- native-hybrid-search-function, runtime-provisioned-per-tenant-search-index, patterns/parallel-retrieval-fusion, metadata-boost-at-query-time, upload-then-poll-indexing, patterns/tool-surface-minimization.
- companies/cloudflare — parent org.