Skip to content

CONCEPT Cited by 8 sources

Retrieval → ranking funnel

Definition

The retrieval → ranking funnel is the canonical two-stage architecture for recommendation, search, and recommendation-like systems at scale:

  1. Retrieval (stage 1). A cheap, high-recall primitive narrows an intractably large candidate population (millions to billions of items) to a rank-tractable set — typically 10² to 10⁴ candidates.
  2. Ranking (stage 2). A more expensive model — cross-encoder, LLM, or a large MTML network — scores or orders the narrowed set and produces a ranked short-list (top-K) to present to the user.

The asymmetric cost structure — retriever runs on every request against a huge pool, ranker runs only over a small narrowed set — is what makes the overall system affordable at production request volume.

SilverTorch face — widened funnel (2026-05-26)

Meta's SilverTorch post (Source: sources/2026-05-26-meta-silvertorch-index-as-model-a-new-retrieval-paradigm-for-recommendation-systems) describes a structural shift in how much intelligence runs at the retrieval stage:

"In traditional service-based systems, retrieval is usually constrained to a relatively narrow ANN result set, scored mostly by simple embedding similarity, with richer relevance modeling deferred to late-stage ranking. SilverTorch unlocked headroom. By keeping ANN search, filtering, and scoring inside one model, it can widen the funnel substantially. Instead of handing only a small set of candidates downstream, it can bring one to two orders of magnitude more candidates through additional learned relevance layers before final ranking. That makes retrieval contribute meaningfully to recommendation quality, not just a fast pruning step."

Two new capabilities run inside the retrieval forward pass under Index as Model:

  • Neural reranking — "multi-layer perceptrons, stacked self-attention, or more structured interaction models such as mixture of logits" applied to retrieval candidates, producing a richer relevance score than dot-product similarity.
  • Multi-task scoring — composite score over multiple engagement-action probabilities (like / share / comment) produced inside retrieval, not deferred to late-stage ranking.

This does not abolish the funnel — late-stage ranking still runs over the narrowed set. What changes is the width of the candidate pool that reaches ranking and the quality of the survivors: "more candidates survive early retrieval, and they are screened by more sophisticated, multi-objective scoring before being passed to the final ranking." The wiki's earlier "retriever recall is the ceiling" property (below) becomes a substantially less binding constraint when retrieval can run multi-task scoring on a 1-2 orders-of-magnitude wider pool.

Structural properties

  • Retriever recall is the ceiling on end-to-end accuracy. If the correct / best item doesn't survive retrieval, no amount of ranking quality can recover it. The Meta Friend Bubbles post (sources/2026-03-18-meta-friend-bubbles-enhancing-social-discovery-on-facebook-reels) states this directly: "By explicitly retrieving friend-interacted content, we expand the top of the funnel to ensure sufficient candidate volume for downstream ranking stages. This is important because, without it, high-quality friend content may never enter the ranking pipeline in the first place."
  • Ranker precision is the ceiling on how cleanly the top-K isolates the right answer. The two ceilings compose multiplicatively — both must meet their bars independently.
  • Expanding top-of-funnel is a dial. When a new candidate class (friend-interacted Reels, a new content vertical, a new index) is missing from the ranker's output, the fix is often at retrieval, not ranking.

The retriever choice space

  • Heuristic retrieval. Domain rules — ownership, social graph + closeness threshold, code-graph traversal, time windows. Fast, interpretable, limited to encoded knowledge. Used by Meta RCA (ownership + code graph) and Meta Friend Bubbles (close-friend candidate sourcing via viewer-friend closeness).
  • Lexical (BM25). Term-frequency scoring. Fast, interpretable, limited to surface keyword match. Canonical for text search.
  • Vector + ANN. Learned embeddings + approximate nearest-neighbour search. Handles semantic similarity; needs embedding infra.
  • Hybrid. Combined lexical + vector. Industry default for document search.
  • Two-tower / multi-tower recall models. Purpose-trained retrieval models for recommendation — typical at Meta / Google / YouTube / TikTok scale. Not named by this Meta post but the standard family for "friend-interacted content retrieval."

Choice is domain-driven. Monorepo RCA has structured ownership + code graph (heuristics win); open-domain document search benefits from hybrid; Reels-scale recommendation typically combines heuristic closeness-based retrieval with embedding-based video-similarity recall.

The ranker choice space

  • MTML models. Multi-task multi-label deep networks with shared encoders + task-specific heads, optimising many engagement targets jointly. The industry default for large-scale recommendation ranking (Meta, Google, TikTok). Meta Friend Bubbles: early-stage + late-stage MTML models with new bubble-conditioned tasks.
  • LLM ranker. LLM scores / orders the narrowed set in natural language. Meta RCA: fine-tuned Llama-2 (7B) running ranking-via-election.
  • Cross-encoder. Smaller Transformer scoring (query, candidate). Cheaper than LLM; less reasoning capacity. Canonical for document search reranking.
  • Pointwise classifier. Small domain-trained model outputs a score per (query, candidate). Cheapest; weakest.

Canonical wiki instances

  • Meta Friend Bubbles (2026-03-18) — recommendation instance. Heuristic closeness-based retrieval + MTML ranking with conditional-probability bubble objective. Canonical datum for expanding top of funnel as the fix for missing candidate class.
  • Meta RCA (2024-06) — RCA / LLM instance. Heuristic ownership + code-graph retrieval + Llama-2 7B ranker via ranking-via-election. See retrieve-then-rank-llm for the LLM-specific pattern.

Relation to other framings

  • retrieve-then-rank-llm is the LLM-specific pattern instance of this funnel concept — when the stage-2 ranker is specifically an LLM.
  • llm-cascade is the sibling cascade pattern at model-size level (small LLM → large LLM), orthogonal to the stage level (retriever → ranker) described here. Both are cascades; they compose.

Caveats

  • Cascading failure modes. A bug in the retriever (missing rule, stale embeddings, broken graph traversal) can systematically bias the candidate set in a way the ranker cannot detect. End-to-end ground-truth evaluation catches this; unit-testing each stage does not.
  • Retriever recall must be measured as a first-class metric — not just ranker precision / NDCG. Without a retriever-recall number, you don't know where your ceiling is.
  • Expanding top-of-funnel is not free. More candidates means more ranker cost. The dial is bounded by ranker latency / throughput budget.
  • Feedback loops bias the retriever. If the retriever only surfaces items the ranker already scored highly, the system can collapse onto a shrinking candidate set. A continuous feedback loop (closed-feedback-loop-ai-features) must feed all candidate sources, not just top-ranked outcomes.

Seen in

  • sources/2026-10-02-databricks-real-time-retail-intelligence-building-e-commerce-recommenda — canonical three-stage e-commerce recsys funnel, run two ways. Databricks' retail reference architecture runs an explicit 3-stage funnel — Stage 1 candidate retrieval (AI Search hybrid ANN + hard metadata filters → 200–500 candidates), Stage 2 feature assembly + LightGBM conversion-probability scoring (CPU, no GPU), Stage 3 deterministic business-rule re-ranking (inventory/proximity/diversity/promo) → truncate to 10–20. The same funnel is executed offline nightly (Path A, precomputed per user into Lakebase) and synchronously at request time (Path B, inside one MLflow PyFunc predict() — co-located inference). Notable: hard eligibility filters are pushed into retrieval, not applied post-scoring, so out-of-stock/ineligible items never enter the candidate set — a recall-shaping move at the funnel's top boundary.
  • sources/2026-03-18-meta-friend-bubbles-enhancing-social-discovery-on-facebook-reels — canonical recommendation-system instance with explicit top-of-funnel expansion.
  • sources/2024-08-23-meta-leveraging-ai-for-efficient-incident-response — canonical LLM-ranker instance; same structural pattern in a non-recommendation domain.
  • sources/2026-02-27-pinterest-bridging-the-gap-online-offline-discrepancy-l1-cvr — Pinterest ads funnel (retrieval → L1 → L2 → auction) is the wiki's canonical multi-stage recommendation funnel example with two explicitly-tracked recall metrics at the L1 boundary: retrieval recall (among final auction winners, how many came from L1 output?) and ranking recall (among top-K by downstream utility, how many appeared in L1 output?). Pinterest's closing observation generalizes: "beyond a certain point, L1 model quality is not the bottleneck — the funnel and utility design are." Canonical demonstration that recall ceilings at each stage boundary cap end-to-end quality independent of within-stage model improvement — a structural lesson beyond two-stage simplifications.
  • sources/2026-09-15-google-bypassing-inference-bottlenecks-accelerating-complex-ai-search-with-retrieve-for-train — query fan-out as a retrieval-stage widening technique. Google Research's Retrieve-for-Train targets set-valued retrieval: decomposing one broad prompt into a coherent, complementary slate of sub-queries ("camping gear" → tent, sleeping bag, stove, headlamp) rather than one best match. Quality is defined by non-decomposable set-level properties (diversity, coverage, complementarity, coherence) that only exist when scoring the whole slate — breaking the traditional pointwise learning-to-rank framing where each retrieved item is scored in isolation. The sub-queries feed a downstream retriever, so fan-out sits at the top of the funnel; the paraphrastic collapse failure mode (redundant near-synonym sub-queries) is the fan-out analogue of a low-recall retrieval stage that caps everything downstream.
  • sources/2026-05-11-yelp-ml-based-ranking-using-nrtsearch — funnel-stage fusion (physical, not logical). Yelp keeps the logical two stages (retrieve candidates in Nrtsearch, then ML-score/re-rank) but collapses them into one process: the ranker runs inside the search node via the Inference Plugin instead of as a downstream service. The two-stage-as-two-services implementation was network-bound at scale (concepts/network-round-trip-cost); co-locating inference removes the hop. Contrast the SilverTorch face above (funnel widened by in-model scoring) — here the funnel stays the same width but the service boundary between retrieval and ranking is dissolved.
  • heuristic-retrieval — the stage-1 primitive used by both canonical instances.
  • llm-based-ranker — one stage-2 option.
  • multi-task-multi-label-ranking — the recommendation-ranker option.
  • concepts/retrieval-ranking-funnel — the document-search stage-1 option.
  • retrieve-then-rank-llm — the LLM-specific pattern form.
  • llm-cascade — the sibling cascade at model-size level.
  • systems/meta-friend-bubbles — recommendation instance.
  • systems/meta-rca-system — LLM/RCA instance.

Merged aliases

  • reranking
  • similarity-tier-retrieval- cross-encoder-reranking
  • hybrid-retrieval-bm25-vectors
  • hybrid-search
  • reciprocal-rank-fusion
Last updated · 766 distilled / 2,225 read