Skip to content

Cloudflare AI Search: give your agents a search engine for your data

Summary

Cloudflare announces developer-experience improvements to AI Search that eliminate the need to manually stitch together Workers AI, AI Gateway, Vectorize, R2, and Browser Run. AI Search now handles crawling, ingestion, embedding, and retrieval automatically. The post also introduces the Cloudflare Dev Stack MCP — a multi-instance search surface that gives coding agents current, cited docs from across the Cloudflare developer ecosystem — and previews predictable pricing where embedding and reranking are free with default Workers AI models.

Key takeaways

  1. Simplified indexing: A single wrangler ai-search instance create command (or --parse-type discover for sites without sitemaps) replaces manual orchestration of multiple primitives. Browser Run's /crawl powers sitemap-less ingestion by following links (Source: article §"Index a collection of data").

  2. Multi-instance namespace search: A single API call fans out across all instances in a namespace and returns merged, cited results tagged by source instance — canonical use of patterns/parallel-retrieval-fusion at the namespace level (Source: article §"Combine the instances into one search").

  3. Two deployment paths — Worker binding or public endpoint: Bind a namespace to a Worker for programmatic control (e.g., an MCP server tool), or enable public URLs on the namespace for zero-code /search and /mcp endpoints (Source: article §"Option A" / §"Option B").

  4. Custom domains + Cloudflare Access: Public endpoints can sit behind custom domains (e.g., search.example.com/mcp) and optionally behind Cloudflare Access for private search instances (Source: article §"Brand it and lock it down").

  5. Bot policy compliance: AI Search's crawler identifies itself as Cloudflare-AI-Search, follows robots.txt, and respects origin bot controls — the same posture as Browser Run (Source: article §"AI Search respects all bot policies").

  6. Predictable pricing model: Embedding and reranking are free when using default Workers AI models. Ingestion billed at $0.75/1M tokens (5M free), storage at $2/GB-month (10 GB free), semantic queries at $0.75/1K (2K free), full-text queries at $0.10/1K (2K free). Chunking uses ~10% overlap, reflected in token billing (Source: article §"Preview pricing").

  7. Example bill: 20M text tokens + 1M image tokens + 30K semantic queries/month ≈ ~$35/month on Workers Paid plan, with embedding and reranking at $0 (Source: article §"Example bill with preview pricing").

  8. Dogfood instances: Cloudflare.com, Developer Docs, and Blog all run on AI Search with hybrid search (semantic + keyword combined) — same infrastructure available to customers (Source: article §"Powering search on our Blog").

Architectural details

  • Namespace → instances → search: A namespace groups multiple instances (one per data source/site). The search() call accepts instance_ids to specify which instances to query. Results come back with instance-level attribution.
  • Hybrid search: Semantic and keyword together in one query — handles both open-ended questions and exact name/keyword lookups.
  • Ingestion chunking: ~10% overlap between chunks (reflected in the × 1.1 billing multiplier in the example bill).
  • Image processing: Add-on at $0.50/1M tokens on top of base ingestion.
  • Storage estimate: ~10 KB/document text, ~1 MB/image.

Operational numbers

Metric Value
Example bill (20K docs + 1K images + 30K queries) ~$35/month
Embedding cost (default models) $0
Reranking cost (default models) $0
Base ingestion $0.75/1M tokens
Semantic query $0.75/1K queries
Full-text query $0.10/1K queries
Storage $2/GB-month
Chunking overlap ~10%

Caveats

  • Pricing is preview — subject to change before billing begins; billing is not yet enabled.
  • The post is primarily a DX improvement announcement; the core search architecture (hybrid BM25 + vector, reranking) was established in the 2026-04-16 AI Search launch.
  • No latency, throughput, or recall numbers disclosed in this post.
  • No detail on how --parse-type discover handles JavaScript-rendered pages or dynamic content beyond "following links."

Source

Last updated · 766 distilled / 2,225 read