Skip to content

CONCEPT Cited by 3 sources

Working-set memory

Definition

Working-set memory — the subset of a database's data + indexes that is actually accessed by the live query workload within a given time window and therefore needs to reside in memory for steady-state performance. Not the total dataset size and not the full index size — the hot slice of each.

A system where working-set > available cache pages in-and-out constantly, which means each query tends toward disk-latency rather than memory-latency. A system where working-set ≤ cache sits in an entirely different performance regime.

Informally: "the dataset that has to fit in RAM for the database to feel fast."

Common working-set shapes

  • Recent-time-window (write-heavy, time-ordered reads): only the last N hours / days of data is hot. Index on timestamp stays small in the active range; cold data is fine on disk. Most analytics / observability / IoT workloads.
  • High-cardinality hot tail (large product catalog, long tail of rarely-accessed items + small head of viral items): the top few % of keys dominate access. Working set is key-count-bounded, not data-size-bounded.
  • Broad low-locality (random reads across the dataset): no hot slice. Working set ≈ entire dataset. The hardest case — capacity planning requires RAM proportional to dataset size.
  • Index-heavy point lookups (unique-field lookups by secondary index): the index itself, not the data pages, is the hot slice. Common in transactional workloads.

Identifying which shape applies drives schema + index strategy.

Working-set vs WiredTiger cache

In MongoDB, "fit in WiredTiger cache" is the operational realization of "working-set fits in memory." The cache is usually ~50 % of system RAM; the OS filesystem cache layers below it; together they form the working-set residence tier.

Empirical diagnostic: observe eviction rate under representative load. Rising evictions under steady workload → working set doesn't fit. Flat evictions + high cache-hit ratio → it does.

MongoDB Cost of Not Knowing Part 3 specifically surfaces the index as the working-set component that didn't fit in appV6R0 (3.13 GB > 1.5 GB cache). Data pages were compressed enough to fit fine; the _id B-tree wasn't.

Levers to fit working set in memory

Roughly in order of first-pass preference:

  1. Shrink the schema (what MongoDB's Cost of Not Knowing series does — field-name shortening, data-type tightening, bucket pattern, computed pattern, dynamic schema).
  2. Right-size the bucket window — wider buckets = fewer documents = smaller index. The appV6R0 → appV6R1 pivot.
  3. Drop unnecessary secondary indexes. Every additional index adds its own B-tree to the working set.
  4. Add RAM. Obvious and sometimes correct, but typically the least-efficient lever — cloud RAM costs 10× a storage byte.
  5. Shard. Splits working set across machines. Expensive in operational complexity; last resort.
  • Working-set vs total dataset. Total dataset is the full on-disk footprint; working set is the hot slice. Capacity planning cares about working set; durability planning cares about total.
  • concepts/cache-locality — the code-level analogue: CPU cache working set. Same principle (hot subset fits in fast tier); different tier.
  • disk-throughput-bottleneck — disk throughput is the symptom when working set doesn't fit. The bottleneck class is usually the signal; working-set sizing is the fix.

Seen in

  • sources/2025-10-09-mongodb-cost-of-not-knowing-mongodb-part-3-appv6r0-to-appv6r4 — appV6R0's 3.13 GB _id index is the working-set component that overflowed the 1.5 GB WiredTiger cache on a 4 GB-RAM test machine, making "memory/cache the limiting factor rather than document size." appV6R1's quarter-bucketing recovered working-set fit by collapsing document count ~3×.
  • — Berquist canonicalises working-set fit as the primary data-size sharding trigger (not disk capacity — "These days, that's not the problem. For example, Backblaze purchased over 40,000 16TB drives!"). "One thing to consider is how large your working set is, and how much of that fits into RAM. As less of your active data fits in memory and more queries need to read from disk, query latency will increase." Connects working-set sizing directly to the scaling-ladder decision of when to climb from vertical scaling to horizontal sharding.

  • sources/2026-08-31-databricks-autoscaling-lakebase-postgres — Lakebase Postgres estimates the current working set online to size its compute cache and drive autoscaling. Because exactly counting every accessed page would cost too much memory, it uses a time-windowed HyperLogLog to approximate distinct-pages-in-the-last-N-minutes, adaptively picks the window (steady workloads plateau; a finished heavy workload shows a jump), and projects growth forward to pre-size the cache (working-set-projection-for-cache-sizing). Canonical wiki example of working-set as a first-class, continuously re-estimated autoscaling signal rather than a static capacity-planning quantity.

  • sources/2026-09-10-databricks-improving-lakebase-postgres-compute-cache — the follow-on: once the working set is estimated, Lakebase must physically fit it in the fastest tier. On large fixed computes it drops the interim NVMe local file cache and sizes shared buffers to 75% of DRAM so the hot working set lives in DRAM rather than cascading to NVMe — driving the compute cache hit rate to ≈100%. Concrete illustration that "fit the working set in memory" is a tiering decision (DRAM vs NVMe vs object storage), not just a total-capacity one.

  • sources/2026-09-30-databricks-a-practical-guide-to-cost-optimization-with-lakebase-postgre — Working set as the primary compute-sizing input, stated in cost terms. Lakebase's cost guide: "The single most important input to sizing is your working set … A 2500 GB database with a 20 GB hot working set does not need 2500 GB of RAM." Up to 75% of a compute's RAM is available as compute cache; when the working set fits, most reads are memory hits with consistent latency, otherwise Postgres fetches missing pages from storage (slower, with latency variability). The guidance: pick a compute whose cache exceeds the working set (plus headroom for query complexity, concurrency, latency targets), and read the Lakebase Metrics dashboard's 5-min / 15-min / 1-hour working-set windows against available cache. A falling cache hit ratio is the leading undersizing indicator — it slips before latency visibly degrades. Pairs with sources/2026-08-31-databricks-autoscaling-lakebase-postgres (working set as a continuously re-estimated autoscaling signal) as the operator-facing sizing guidance.

Last updated · 766 distilled / 2,225 read