Skip to content

SYSTEM Cited by 1 source

MLeap

MLeap is an open-source model serialization + execution format and runtime that lets ML models trained in frameworks like Spark ML, scikit-learn, TensorFlow, and XGBoost be exported as a portable bundle and executed at serving time on the JVM without the training framework's runtime — a low-latency inference path for production scoring. It sits on the serving side of the training-serving boundary: train once, export an MLeap bundle, load it into a serving process.

Why it matters for system design

  • Serving-side portability. An MLeap bundle is a single deployable artifact that a serving process loads and runs, decoupling inference from the (heavier) training stack.
  • JVM-native inference enables co-location. Because MLeap runs in-process on the JVM, models can be executed inside a Java serving node rather than behind a separate scoring service — the enabling substrate for co-located inference.

Seen in

  • Yelp — ML based ranking using Nrtsearch (2026-05-11). MLeap is the inference format for Yelp's ML platform (alongside MLflow as the model store). The Nrtsearch Inference Plugin's ML Scorer runs MLeap-based inference per model, in-JVM on search-replica nodes, over XGBoost and neural models. Yelp names supporting inference engines beyond MLeap as future work.

Stub page. Expand as dedicated sources arrive.

Last updated · 766 distilled / 2,225 read