PATTERN Cited by 3 sources
Co-located inference in the serving layer¶
Definition¶
Co-located inference in the serving layer runs the ML scoring/ranking model inside the same process (or node) that holds the data being scored — typically the search-index or retrieval layer — instead of shipping candidates and their features over the network to a separate scoring service. Feature extraction and inference happen where the candidate documents already live, so the model reads features straight off the index document (or the request), scores in-process, and returns a ranked result set. The pattern collapses the classic two-hop retrieve → (fetch features) → score elsewhere flow into one hop.
It is the serving-layer counterpart of "move compute to the data": rather than moving large candidate sets and fat feature payloads to the model, move the model to the data.
Problem it solves¶
The two-stage retrieval → ranking funnel is often implemented as two services: a retrieval service returns candidate IDs, a separate inference service scores them. That decoupling is clean but becomes network-bound when either dimension grows:
- Large candidate sets — thousands+ of eligible candidates must all cross the wire to be scored.
- Large per-candidate feature payloads — rich document/user features multiply the bytes shipped per candidate.
Both inflate network transfer and serialization cost, which come to dominate end-to-end latency — the actual model math may be cheap by comparison. Co-locating inference removes the hop entirely.
Mechanism¶
- Embed the model in the serving node. The retrieval/index node hosts an in-process model server (same JVM/runtime), loading model bundles at bootstrap from a model registry.
- Extract features locally. Features are read from the index document or the incoming request — not fetched from a remote feature store at score time. This is the load-bearing constraint (see below).
- Score in-process, return ranked results. The node retrieves candidates, scores each with the model(s), aggregates a final score, and returns the ranked set — one round trip from the client's perspective.
The load-bearing constraint¶
The pattern's power comes from a real restriction: every feature the model needs must be derivable from data already co-located (the index document or the request). If a feature requires a live external join — real-time user-session state, a value from another service — it either has to be denormalized into the index at write time, passed on the request, or the candidate can't be scored in-process. Adopters accept this constraint deliberately; it is the price of removing the network hop.
Seen in¶
-
Databricks — real-time retail recommendations (2026-10-02). The real-time (Path B) recommendation endpoint is a custom MLflow PyFunc model that "orchestrates the multi-stage pipeline internally — querying AI Search, performing Lakebase lookups, running LightGBM inference, and applying business rules within a single predict() call." Rather than a retrieval service → feature service → scoring service → re-rank service chain (several network hops), the whole funnel executes inside one endpoint process. A slightly different flavor from the Yelp/Meta instances (which embed the model in the search node): here the serving endpoint is the orchestrator that pulls from retrieval + feature stores and scores in-process, collapsing the multi-service funnel into one request-response cycle to hit a low-2-digit-ms budget. The embedding-blend step (long-term ⊕ session) is likewise "a lightweight computation performed within the endpoint, not a separate model call." Explicit latency- budget fallback to pre-computed Path A results on overrun.
-
Yelp — ML based ranking using Nrtsearch (2026-05-11). The canonical instance. Yelp's Nrtsearch Inference Plugin embeds an in-JVM, TensorFlow-based model server on replica nodes, extracts features from the index document / request, and runs MLeap-based scoring (XGBoost + neural models) inside the Lucene Function Score query or rescoring phase — explicitly to eliminate the standalone inference service and "reduce latency and serialization overhead." Mandates that "all features be extracted from the index document or request parameters." Two tiers: custom Java Scorer module, or a generic scorer defined via Lucene expressions.
- Meta — SilverTorch: Index as Model (2026-05-26). The GPU-native extreme of the pattern: item index, eligibility filter, scoring layer, and user tower are all "a tensor or operator inside a single PyTorch model" — "one artifact to deploy, one forward pass to run." The filter result "is already inside the model, [so] it can flow directly into ANN search without a separate service call." Widens the funnel by running multi-task scoring in-graph rather than deferring it to a downstream service.
Consequences / tradeoffs¶
- Wins: removes a network hop and its serialization cost; lets the funnel widen (score more candidates affordably); one deploy artifact; features are read at native speed off the index.
- Costs: the feature-locality constraint restricts which models are expressible; couples model deployment to the search fleet's lifecycle (a scorer change may force a rolling restart — see patterns/staged-rollout); the serving node now carries model-server memory/CPU (or GPU) alongside retrieval, changing its resource profile; model-quality monitoring must be added separately from system-health monitoring.
- Rollout shape: because scoring lives in the data plane, model/scorer changes ride the serving fleet's deploy path — canary node → monitored → rolling restart (Yelp), one-node-first model updates without restart for pure model swaps.
Related¶
- concepts/retrieval-ranking-funnel — the two-stage architecture this fuses
- concepts/network-round-trip-cost — the cost removed
- concepts/network-bound-vs-compute-bound — the regime where this wins
- concepts/feature-store — what the two-stage flow fanned out to
- systems/nrtsearch — Yelp's in-search-engine instance
- systems/silvertorch — Meta's in-model (Index as Model) instance
- patterns/staged-rollout — how co-located models get deployed safely