Skip to content

CONCEPT Cited by 8 sources

Feature store

Definition

A feature store is the class of ML-infrastructure systems that manage and deliver feature data — the numerical/categorical signals a model consumes at both training time and inference time — as a first-class shared substrate across teams.

Concretely a feature store typically provides:

  • Feature definitions — a declarative spec for each feature (name, type, source, transformation, freshness SLO).
  • Offline store — historical values for training (often cheap object storage + a table format).
  • Online store — low-latency lookup at inference time (often a KV store like DynamoDB, Redis, or a DynamoDB-compatible internal store like Dynovault).
  • Ingestion — batch + streaming + direct-write pipes from source systems into both stores.
  • Serving API — runtime interface model-serving code calls to fetch features for a request.

Why feature stores exist as a separate thing

Models don't ship one feature at a time. A ranker typically wants dozens of features per candidate × hundreds of candidates per query — thousands of lookups. Having every ML team build their own ingestion, store, and serving gets expensive and inconsistent fast. A shared feature store:

  1. Decouples feature engineering from model serving — ML engineers write transformations; serving infrastructure is abstracted away.
  2. Keeps training and serving on the same feature values — the canonical failure mode without a feature store is training/serving skew (a feature computed one way at training time and a slightly different way at serving time). Shared definitions + offline/online stores that share a lineage fix this; adjacent to concepts/training-serving-boundary but not the same axis — this is feature-data consistency, not compute- fleet unification.
  3. Amortizes ingestion across many models — one pipeline feeds the features; N models consume.

Landscape

Named in the Dropbox post as evaluated options:

  • Feast — open-source.
  • Hopsworks — open-source with managed offering.
  • Featureform — open-source.
  • Feathr — open-source (LinkedIn origin).
  • Databricks Feature Store — platform-native.
  • Tecton — commercial.

Dropbox chose Feast for its clean definitions/infra separation and adapter ecosystem, then layered their own serving tier (Go + Dynovault) behind it — a common pattern for large orgs whose on-prem or bespoke infrastructure doesn't match any vendor's shape. (Source: sources/2025-12-18-dropbox-feature-store-powering-real-time-ai-dash)

Load-bearing design axes

When evaluating or building a feature store, the knobs that dominate:

  • Latency budget — sub-100ms is typical for ranking; sub-10ms is aggressive but occasionally needed for search autocomplete.
  • Fan-out shape — per-query how many feature lookups? Per-user? Per-candidate? Amplifies all other axes.
  • Freshness requirement — some features can be days stale (content embeddings); some must be seconds-fresh (recent-interaction signals).
  • Ingestion shape — does data arrive in bulk (batch), as a stream (CDC, event), or directly written by another model (precomputed scores)? See hybrid-batch-streaming-ingestion.
  • Training/serving consistency — same feature values available in both paths, same semantics.

Seen in

  • sources/2026-10-02-databricks-real-time-retail-intelligence-building-e-commerce-recommenda — feature store as the training-serving-consistency layer in a recsys. Databricks' retail reference architecture uses the Databricks Feature Store to manage both offline features (for training) and online features (for serving), "ensuring training-serving consistency — the same feature definitions used during model training are automatically available at inference time via Lakebase online tables." At serving time (Path B), candidate item IDs trigger point lookups against Lakebase online tables for pre-computed user features (behavioral aggregates, demographic segment, price sensitivity) and item features (conversion rate, trending score, days since listing), combined with request-derived real-time context. A clean instance of the author-once/serve-everywhere property feeding funnel Stage 2 scoring; see concepts/training-serving-boundary.

  • sources/2026-08-17-databricks-how-databricks-feature-store-serves-features-with-sub-second-freshness — Databricks Feature Store's streaming/online path, the platform-native realization named earlier in the Landscape. Adds the author-once, serve-everywhere property (one definition compiles to both offline batch and online streaming), a fresh online path (Kafka → Spark RTM → Lakebase → Model Serving) hitting 200 ms end-to-end p99, and the rolling-window model as the freshness primitive. Every design axis on this page is load-bearing: latency (200 ms), fan-out (per-event rolling upserts), freshness (sub-second via rolling windows), ingestion shape (streaming from Kafka + offline copy for batch), training/serving consistency (offline Kafka copy + point-in-time joins, see backfill-online-features-from-offline-copy). Confirms the "co-located online feature store" observation from the Lakebase tap-pay source at a second, streaming-first realization.

  • sources/2026-07-16-databricks-what-happens-in-the-milliseconds-after-you-tap-pay — Lakebase Postgres as an inline feature store. The fraud model's predict() does its own Lakebase lookup at inference time: extracts card BIN, queries customer_features table, assembles a 12-feature vector, runs CatBoost. Feature lookup takes 8.9 ms p50 — the dominant cost in model time. Structurally this is a co-located online feature store (same database serves both ML features and application profile data), eliminating the need for a separate KV store. Trade-off: no offline/online sync boundary — training features must also come from the same table or be replicated separately.

  • sources/2025-12-18-dropbox-feature-store-powering-real-time-ai-dash — canonical in-wiki introduction; Dropbox's Dash feature-store stack as the worked example, including landscape citation and in-house hybrid-build rationale.

  • sources/2026-01-06-expedia-powering-vector-embedding-capabilities — the feature-store online/offline duality transplanted to an embedding-platform context: online store = vector DB for interactive similarity search, offline store = historical dataset repository for analytics / experimentation / training / backup, with an explicit restore path from offline → online gated on creation date / time range / SQL. Expedia layers systems/feast on top not to register feature views but embedding collections — evidence that the feature-store design pattern (definitions registry + online/offline split + orchestrated ingestion) is a shape, not a schema, and carries over cleanly to vectors.
  • sources/2026-01-06-lyft-feature-store-architecture-optimization-and-evolution — the second canonical major-tech-co instance on the wiki (after Dropbox Dash). Lyft's Feature Store is framed as a "platform of platforms" with the same three structural lanes (batch / streaming / direct-CRUD) plus a unified online-serving layer (dsfeatures) that wraps DynamoDB + ValKey (write- through LRU) + OpenSearch (embeddings) behind a single SDK with full CRUD. Confirms that the five-dimension design axis list on this page (latency, fan-out, freshness, ingestion shape, training/serving consistency) survives a second real-world realisation — every axis is load-bearing in Lyft's post as well. Adds two new primitives the Dropbox post didn't make explicit: config-driven DAG generation for batch feature onboarding and wrapper- over-heterogeneous-stores as the unified-surface implementation shape.
  • sources/2026-06-10-atlassian-architecting-scalable-ml-platforms — Atlassian ML Studio integrates a "central feature store" to reuse trusted features across teams, as part of its cross-functional ML ecosystem integration layer. Mentioned at feature-list level; no architectural detail on the feature store's own internals.
  • sources/2026-06-19-netflix-thinking-fast-slow-for-a-personalized-notification-system — Netflix uses a feature store as the asynchronous communication bridge between Slow (strategic) and Fast (tactical) notification policies (see feature-store-as-policy-bridge). The Slow Policy writes personalized weekly pacing plans; the Fast Policy reads them at send time. Demonstrates feature stores serving not just ML-feature-serving but inter-policy state transfer in hierarchical decision systems.

Merged aliases

  • feature-store-freshness
  • online-vs-offline-feature-store
  • sketching-feature-store- online-feature-serving- feature-freshness
Last updated · 766 distilled / 2,225 read