SYSTEM Cited by 1 source
MLeap¶
MLeap is an open-source model serialization + execution format and runtime that lets ML models trained in frameworks like Spark ML, scikit-learn, TensorFlow, and XGBoost be exported as a portable bundle and executed at serving time on the JVM without the training framework's runtime — a low-latency inference path for production scoring. It sits on the serving side of the training-serving boundary: train once, export an MLeap bundle, load it into a serving process.
Why it matters for system design¶
- Serving-side portability. An MLeap bundle is a single deployable artifact that a serving process loads and runs, decoupling inference from the (heavier) training stack.
- JVM-native inference enables co-location. Because MLeap runs in-process on the JVM, models can be executed inside a Java serving node rather than behind a separate scoring service — the enabling substrate for co-located inference.
Seen in¶
- Yelp — ML based ranking using Nrtsearch (2026-05-11). MLeap is the inference format for Yelp's ML platform (alongside MLflow as the model store). The Nrtsearch Inference Plugin's ML Scorer runs MLeap-based inference per model, in-JVM on search-replica nodes, over XGBoost and neural models. Yelp names supporting inference engines beyond MLeap as future work.
Stub page. Expand as dedicated sources arrive.
Related¶
- systems/mlflow — model store / registry paired with MLeap at Yelp
- systems/xgboost — a model type served via MLeap bundles
- systems/nrtsearch — runs MLeap inference in-process in the Inference Plugin
- concepts/training-serving-boundary — MLeap lives on the serving side
- patterns/co-located-inference-in-serving-layer