Yelp — ML based ranking using Nrtsearch¶
Summary¶
Yelp extended Nrtsearch — its open-source Lucene-based search engine — with an Inference Plugin that embeds ML-based ranking directly inside the search layer, eliminating a previously-standalone scoring service. The original flow was two-stage and network-bound: the web server queried Nrtsearch for candidate IDs, fetched their features from feature stores, then shipped candidates + features over the network to a separate inference service to score and re-rank. This broke down at scale when either the candidate count or the per-candidate feature payload grew large — network transfer and serialization dominated latency. The Inference Plugin collapses the two stages: Nrtsearch retrieves candidates, extracts features from the index document (or request params) in-process, and scores each document with an in-JVM model server — all in milliseconds. The plugin powers Search, Ads, and Home Feed at Yelp, integrates with Yelp's ML platform (MLflow + MLeap), and supports XGBoost and neural models. A hard design constraint is that all features must be derivable from the index document or the request — the price of co-location.
Key takeaways¶
- The motivation is a network / serialization bottleneck, not a model-quality gap. Lucene BM25 scoring is term-based with fixed hyperparameters — no learned weights, no nonlinear feature interactions, no rich user/document signals. The first fix was a two-stage flow: Nrtsearch retrieves candidates → fetch features → standalone inference service scores/re-ranks. It "worked for a lot of use cases" but hit two scaling walls: (1) the eligible-candidate count could be large; (2) the per-candidate feature payload could be large, both leading to "performance bottlenecks due to network transfer overhead and large model feature payloads." (Source)
- The Inference Plugin co-locates feature storage + inference with retrieval. Client → Nrtsearch cluster directly; the cluster (a) searches the index for eligible doc IDs, (b) calls the plugin to extract features, (c) scores each doc with the ML model. Co-locating feature storage and inference "reduce[d] latency and serialization overhead." (Source) This is the canonical wiki instance of co-located inference in the serving layer.
- The load-bearing constraint: features must come from the index document or request parameters. Yelp explicitly mandated that "all features be extracted from the index document or request parameters." This is the tradeoff that makes co-location possible — no fan-out to external feature stores at score time — and it shapes what ranking models are expressible in-plugin.
- Replica nodes host the model server; the primary indexes. Nrtsearch clusters run a built-in TensorFlow-based model server inside the same JVM on replica nodes (which serve search); the primary handles indexing only. On bootstrap, an ML-configured replica loads all model bundles named in its config from MLflow, then serves ML-ranking requests. (Source)
- The ML Scorer is a document-level orchestrator. Per document it: chooses which model(s) to apply, extracts + transforms features via model configs, runs MLeap-based inference per model, and aggregates a final score. Supports multi-model execution, batch inference, safe field access, and error handling. The plugin is callable both in the Nrtsearch Function Score query and the rescoring phase — full flexibility for complex inference.
- Two customization tiers trade velocity for control. Teams either (a) implement the plugin's Java Scorer interface as a custom module (fine-tune every scoring step — for complex needs), or (b) use the built-in generic ML scorer, defining features as Lucene expressions in the query with no module code (fast onboarding, higher developer velocity). This is the escape-hatch shape: a golden generic path plus a code-level override.
- Model updates and scorer changes deploy differently. A pure model update (no scorer change) uses a Jenkins job that hits a Nrtsearch endpoint to update the model on one node first, verifies no errors, then rolls out to the rest — no cluster restart. A scorer change goes to a Canary node first (monitored), then triggers a rolling restart of the cluster, deliberately restarting a small number of nodes at a time to bound the blast radius of undiscovered bugs. (staged rollout)
- Model-quality monitoring is separate from system-health monitoring. Beyond standard health dashboards, Yelp exposes Prometheus metrics (latencies, error rates, model behavior) to alert when a model's performance degrades over time — a model can be "unhealthy" while the system is nominally healthy.
- Future work: GPU-based ranking + multiple inference engines. Yelp plans GPU acceleration for larger/more complex neural rankers, and support for inference engines beyond the current MLeap one.
Systems / concepts / patterns extracted¶
- Systems: Nrtsearch (Inference Plugin, in-JVM model server, Function Score query, rescoring phase), Lucene (BM25 baseline; Lucene expressions for the generic scorer), MLeap (model serialization/serving format), MLflow (model store / registry), XGBoost (a supported model type), BM25 (the term-based baseline being augmented), Yelp ML platform (MLflow + MLeap), Prometheus (model-performance monitoring), Jenkins (model-update rollout).
- Concepts: retrieval → ranking funnel (retrieve candidates, then score/re-rank — here both stages fused into one process), network round-trip cost / network-bound vs compute-bound (the bottleneck the plugin removes), feature store (what the two-stage flow fanned out to, now replaced by index-doc/request feature extraction), training-serving boundary (MLflow/MLeap bundle authored once, loaded into serving).
- Patterns: co-located inference in the serving layer (the headline pattern), staged rollout (canary node → rolling restart, one-node-first model updates).
Operational numbers / specifics¶
- Scoring "happens in a matter of milliseconds to serve live responses" (no latency percentiles, QPS, or fleet-size disclosed).
- Supported aggregation of scores for multi-model requests (final rescore attached to the response object).
- Model server is TensorFlow-based, runs in the same JVM as the search replica; inference itself is MLeap-based per model.
- Generic scorer feature definitions use Lucene expressions (JS-subset expression language) — no Java module required.
Caveats¶
- No quantitative disclosure: no latency numbers (before/after), no candidate counts, no feature-payload sizes, no QPS, no fleet size. The scaling argument is qualitative ("could be large" → "network transfer overhead").
- The index-document/request-only feature constraint is a real limitation, not just a design nicety: features requiring external real-time joins (e.g. live user-session state not on the request) don't fit the co-located model cleanly. The post frames this as "mandating" the constraint rather than as a cost.
- "TensorFlow-based model server" + "MLeap-based inference" + "XGBoost/neural" are stated but the interplay (does MLeap wrap TF? are XGBoost models MLeap bundles?) isn't fully specified. Yelp's ML platform uses MLflow + MLeap; the in-JVM server is described as TensorFlow-based. Recorded as stated.
- GPU ranking and multi-engine support are roadmap, not shipped.
Source¶
- Original: https://engineeringblog.yelp.com/2026/05/ml-ranking-with-nrtsearch.html
- Raw markdown:
raw/yelp/2026-05-11-ml-based-ranking-using-nrtsearch-4bd4fda0.md
Related¶
- systems/nrtsearch — the search engine extended with the Inference Plugin
- systems/lucene — BM25 baseline + Lucene expressions
- systems/mleap — inference format used by the ML Scorer
- systems/mlflow — model store for the ML platform
- systems/xgboost — a supported model type
- concepts/retrieval-ranking-funnel — the two-stage architecture fused here
- concepts/network-round-trip-cost — the bottleneck the plugin removes
- patterns/co-located-inference-in-serving-layer — the headline pattern
- companies/yelp