Skip to content

CLOUDFLARE 2026-10-01

Read original ↗

AI Search is now generally available

Summary

Cloudflare announced the general availability of AI Search (formerly AutoRAG) — its managed index-and-retrieval pipeline stitching Workers AI, Vectorize, R2, and Browser Run into one product. The headline additions on top of the prior hybrid-search release are multimodal retrieval: AI Search now embeds image pixels directly for visual retrieval (via the Qwen3-VL-Embedding model) instead of only captioning-then-embedding, keeping richer embeddings efficient with Matryoshka Representation Learning (MRL); OCR for scanned PDFs; and larger files (text + PDF up to 10 MiB, up from 4 MiB). GA also brings usage-based pricing — billing starts 2026-11-01, structured so a bill is estimable before indexing a single file (ingestion tokens + stored GB + queries, with parsing/chunking/embedding/keyword-indexing/reranking included), on top of a generous free tier on every Workers plan. The article also sketches the query lifecycle (rewrite → embed → parallel vector + keyword search → fuse → optional rerank → return chunks or generate an answer) and previews a video/audio ingestion pipeline and a keyword-engine refactor for larger stores.

Key takeaways

  1. Native image embeddings replace the caption-only hack. The prior implementation was "naive: we would perform object detection, generate a caption, and then embed that text" — images were searchable only through what the caption captured. GA now does both: caption-based textual understanding and native image retrieval that preserves visual detail (texture, layout, composition, spatial relationships, palette) a caption would compress away. (Source: this article)
  2. Qwen3-VL-Embedding is the native-multimodal model. At query time AI Search checks whether the instance's embedding model supports images; if so, a query image is embedded directly into the same vector space as indexed images and text — enabling image→image, text→image, or combined queries ("a bird with similar markings"). (Source: this article)
  3. Graceful degradation for text-only embedding models. If the configured embedding model is text-only, a query image is converted to text via ToMarkdown and searched by the resulting caption — "this gives every model basic multimodal support, while models with native image support get the full visual signal." (Source: this article)
  4. Matryoshka Representation Learning keeps multimodal embeddings cheap. AI Search "leverages Matryoshka Representation Learning (MRL), allowing smaller embeddings to retain useful information while keeping storage manageable and search fast." MRL (arxiv 2205.13147) packs coarse-to-fine information into prefixes of one vector so it can be truncated without re-embedding. (Source: this article)
  5. OCR for scanned PDFs + 10 MiB files. Many PDFs are really scanned images with no extractable text; turning on OCR makes AI Search read each page's text before chunking and embedding. Text files (Markdown/HTML/CSV/JSON) and PDFs now go up to 10 MiB (from 4 MiB). OCR is billed as image-processing ingestion tokens. (Source: this article)
  6. The query lifecycle is now stated explicitly. A query is optionally rewritten, then embedded (directly by a multimodal model, or captioned first by a text-only one); vector and keyword search run in parallel; results are fused and optionally reranked; top chunks are returned, or passed to a generation model to write an answer. This is the canonical retrieve → fuse → rerank / parallel-retrieval-fusion shape. (Source: this article)
  7. Predictable usage-based pricing, billing 2026-11-01. You pay for three things — content ingested, data stored, queries run — and "the work in between (parsing, chunking, embedding with Workers AI models, keyword indexing, and reranking) is included." No instance hours, capacity units, or monthly minimums. Switching a text-only embedding model to a multimodal one doesn't change ingestion cost unless you also process images (add-on fee). (Source: this article)
  8. One tweak vs the August preview pricing: the free tier now grants 1,000 semantic + 1,000 full-text queries (two separate pools) instead of a shared pool of 2,000. (Source: this article)
  9. What's next: a video + audio ingestion pipeline for rich media search; a keyword-search-engine refactor that "scales better… particularly for when you have big data stores where the current implementation has limits"; and simpler ways to create indexes for sites already on Cloudflare so agents can discover/consume content. (Source: this article)

Pricing table (GA)

Category Price Free monthly allotment
Base ingestion $0.75 / 1M tokens 5M tokens †
Image processing (add-on) +$0.50 / 1M tokens (shares 5M pool)
Stored data $2.00 / GB-month 10 GB
Semantic query (hybrid + vector) $0.75 / 1K queries 1,000 queries
Full-text query $0.10 / 1K queries 1,000 queries
Embedding + reranking Free with select Workers AI models (third-party billed separately) N/A

† A single 5M-token ingestion pool per month covers any supported file type (text, images, …). Billing goes live 2026-11-01, with a reminder email first.

Systems / concepts / patterns extracted

  • Systems: AI Search (the product going GA), Workers AI (embedding + reranking models, incl. Qwen3-VL-Embedding), Vectorize (vector index substrate holding the multimodal vectors), R2 (storage), Browser Run (crawl + the ToMarkdown conversion for text-only image queries). Named model: Qwen3-VL-Embedding.
  • Concepts: hybrid search, vector embedding (native image embedding vs caption), Matryoshka Representation Learning (new page), the retrieval → ranking funnel, RAG, multimodal content understanding (caption + native image signals).
  • Patterns: parallel-retrieval-fusion (vector + keyword in parallel, fuse, optional rerank).
  • Prose-only (not minted): ocr-for-scanned-pdf-ingestion, caption-plus-native-image-dual-signal, query-image-to-markdown-fallback, predictable-usage-based-pricing-before-indexing — article-specific implementation/pricing details, recorded as tags here, mapped into the pages above. OCR and the Qwen3-VL model are documented on their system pages, not as concepts.

Operational numbers

  • File-size limit: 10 MiB (text + PDF), up from 4 MiB.
  • Native-multimodal model: Qwen3-VL-Embedding.
  • Billing start: 2026-11-01; preview pricing announced August 2026 (Agents Week).
  • Free tier: 5M ingestion tokens, 10 GB storage, 1,000 semantic + 1,000 full-text queries per month, on all Workers plans.
  • Prices: ingestion $0.75 / 1M tokens; image add-on +$0.50 / 1M tokens; storage $2.00 / GB-month; semantic $0.75 / 1K; full-text $0.10 / 1K.

Caveats

  • GA announcement + pricing post; still light on hard retrieval metrics — no published recall/precision deltas for native image embeddings vs the caption-only baseline, no latency/throughput numbers for the query pipeline.
  • MRL is named but AI Search does not disclose the truncation dimensions used or the recall/size tradeoff curve.
  • Only one native-multimodal embedding model is named (Qwen3-VL-Embedding); pluggability and the full supported-model matrix live in the (external) docs.
  • Video/audio ingestion and the keyword-engine rescale are roadmap, not shipped; the current keyword engine is acknowledged to "have limits" on big stores.
  • The embedding model powering the text/vector half (beyond the multimodal case) is still not named in the post.

Source

Last updated · 766 distilled / 2,225 read