Skip to content

DATABRICKS 2026-10-02

Read original ↗

Real-Time Retail Intelligence: Building E-Commerce Recommendations with Lakebase and AI Search on Databricks

Summary

A Databricks reference architecture for a production e-commerce recommendation system, drawn from a real deployment for a fashion e-commerce platform in Asia (1M+ monthly active users, 100,000+ SKUs, ~1,000 clickstream events/sec). The whole stack — ingestion, feature engineering, training, vector retrieval, and low-latency serving — runs on Databricks and is governed end-to-end by Unity Catalog. The defining design choice is a two-path serving architecture: Path A pre-computes top-N recommendations nightly for every active user and serves them as a plain key-value lookup from Lakebase online tables (single-digit ms); Path B runs the full three-stage retrieval → ranking funnel synchronously inside a single MLflow PyFunc endpoint for session-aware surfaces where intent only exists at request time. In-session signals are passed in the request payload and deliberately bypass the lakehouse to avoid ingestion latency on the hot path.

Key takeaways

  • Two serving paths, split by whether the context is known in advance. Path A (homepage carousels, category rankings, email/push) is pre-computed because user + surface are known ahead of time; Path B (similar-items, complete-the-look, re-ranked search) is computed live because context "only exists at request time." Both target low 2-digit-ms responses. (Source: this article)
  • Path A is an async-projected read model. A nightly Databricks Workflows job runs the full funnel offline for every active user and writes a top-N list (50–100 items per surface) into Lakebase online tables keyed by user_id + surface. Serving is "no model inference, no vector search, no feature assembly — just a direct read." (Source: this article)
  • Path B co-locates the whole funnel in one predict() call. The endpoint is a custom MLflow PyFunc model that orchestrates all three stages — query AI Search, point-lookup Lakebase, run LightGBM, apply business rules — "within a single predict() call," collapsing what would otherwise be several network hops. (Source: this article) — an instance of co-located inference.
  • Hard filters are pushed into retrieval, not applied after scoring. AI Search does hybrid retrieval — ANN similarity plus hard metadata filters (category eligibility, regional inventory, min-stock) — so "items that are out of stock or ineligible never enter the candidate set." Returns 200–500 candidates in a single round-trip. (Source: this article)
  • The scoring model is gradient-boosted trees, not a GPU model. LightGBM predicts conversion probability per user-candidate pair "in a single batch inference call without GPU resources." Cross features (user-category affinity, brand overlap) are built inside the endpoint. (Source: this article)
  • Blended query embedding fuses long-term + in-session intent. For session-aware surfaces the endpoint combines the user's pre-computed long-term preference embedding (from Lakebase) with a session embedding from recently viewed items, weighted toward recent signals — "a lightweight computation performed within the endpoint, not a separate model call." (Source: this article)
  • Business rules are config-driven and read at serving time. Promotional weights, diversity thresholds, and inventory cutoffs live in a managed configuration table, "allowing commercial teams to adjust rules without redeploying the model." (Source: this article)
  • Explicit fallback for resilience. If the real-time path exceeds its latency budget, the system "gracefully degrades to serving cached popular items or the user's pre-computed recommendations from Path A" — i.e., Path B falls back to Path A. (Source: this article)
  • Cold start handled on both axes. New users get a default embedding from demographics (location, device, sign-up context), placing them in a behavioral cluster that converges as they interact; new SKUs get an embedding from attributes + product-image visual features, inheriting scores from nearest neighbors, surfacing at the next daily batch. (Source: this article)
  • Training-serving consistency via the feature store. The Databricks Feature Store manages both offline (training) and online (serving) features so "the same feature definitions used during model training are automatically available at inference time via Lakebase online tables." (Source: this article)
  • Weekly champion/challenger retraining with drift detection. Models retrain weekly (MLflow tracking/versioning); new versions deploy alongside production with traffic gradually shifted on online metrics. Automated drift detection on feature/score distributions triggers investigation or accelerated retraining; serving logs are joined back to training via request-level IDs for leakage-free feedback, with position-aware training to counter display-position bias. (Source: this article)

Architecture

Ingestion

~1,000 events/sec (views, searches, add-to-cart, purchases, session metadata) land in Unity Catalog Delta tables via Zerobus Ingest (Lakeflow Connect) — no self-managed broker. Zerobus accepts any standard Kafka producer client (Java/Python/Go) by repointing the bootstrap server; teams already on Kafka can use Structured Streaming + Declarative Pipelines instead. Crucially, clickstream flows to the lakehouse only for offline feature computation/training; in-session signals on the inference hot path (Path B) are sent in the API request payload and bypass lakehouse storage entirely.

Medallion layering

Bronze/Silver/Gold: - Bronze — raw append-only event streams + reference data (user behavioral/ transactional/preference/demographic signals; item catalog/performance/visual/ freshness signals; contextual temporal/geographic/environmental/session signals). - Silver — sessionized behavior, cleaned product features, aggregated engagement signals (7-day view counts, trending scores), user-product interaction matrices. - Gold — model-ready user/item feature vectors, labeled training datasets, and pre-computed recommendation lists.

Refresh cadences differ: behavioral aggregates refresh daily via Databricks Workflows, full catalog syncs weekly, user+item embeddings recompute daily.

Serving (two paths)

Path A (pre-computed, known user+surface):
  nightly Workflows job → full funnel offline per user
    → retrieve user embedding → batch ANN on AI Search
    → score with LightGBM (Gold features) → business rules
    → write top-N (50-100/surface) to Lakebase online tables
  serving: app → KV lookup (user_id + surface) → ranked list   [<2-digit ms]

Path B (real-time, session-aware):
  app → REST → MLflow PyFunc endpoint:
    Stage 1  blended query embedding (Lakebase long-term ⊕ session)
             → AI Search hybrid retrieval (ANN + hard filters)
             → 200-500 candidates (1 round-trip)
    Stage 2  Lakebase point-lookups (user+item features) + request-derived
             context → feature vectors → LightGBM batch score (no GPU)
    Stage 3  re-rank: inventory/proximity/diversity/promo → truncate 10-20
  fallback: latency budget exceeded → cached popular items or Path A list

Model lifecycle

Weekly retrain via Workflows; MLflow experiment tracking/versioning; champion/challenger with gradual traffic shift. Metrics: AUC-ROC, NDCG@K, Recall@K, Log Loss, plus business KPIs (CTR, conversion, revenue/session). Automated drift detection; request-level serving↔training correlation; position-aware training.

Operational numbers

  • 1M+ monthly active users; 100,000+ SKUs.
  • ~1,000 ingestion events/sec.
  • Path A pre-computes 50–100 items per surface per active user, nightly.
  • Path B: AI Search returns 200–500 candidates per request in one round-trip; final list truncated to 10–20 items.
  • Industry-benchmark framing: effective personalization lifts conversion 10–30% (not a measured result of this deployment).
  • Both paths target low 2-digit-ms personalized responses.

Caveats

  • Tier-3 Databricks blog with vendor framing; this is a reference architecture distilled from one unnamed customer deployment, not an independently reported production incident/benchmark.
  • Concrete latency figures are given qualitatively ("<2-digit ms", "single round-trip") rather than as measured P50/P99 distributions; the 10–30% uplift is an industry benchmark, not this system's result.
  • Path A freshness is bounded by the nightly batch (previous-day signals/ inventory); acceptable because "long-term preferences and brand affinities evolve over days, not minutes," but a known staleness trade-off.
  • "Databricks AI Search" is Databricks' managed vector-search product (Mosaic AI Vector Search); the article uses the newer "AI Search" branding.

Source

Last updated · 766 distilled / 2,225 read