Skip to content

CONCEPT Cited by 3 sources

Predicate pushdown

Predicate pushdown is the query-optimization technique of pushing filter predicates (WHERE clauses) down into the storage layer so that irrelevant data is never read into the execution engine. In columnar formats like Parquet, this means consulting per-row-group min/max statistics to skip entire row groups that cannot satisfy the predicate.

Effectiveness depends on file size

Row group statistics work best when row groups are substantial. Tiny Parquet files (e.g. 500 KB written every 30 seconds from a streaming flush) often have row groups too small to prune meaningfully — the statistics span so little data that most predicates match anyway. Larger files (32+ MB) with well-populated row groups enable significant data-skipping savings (Source: sources/2026-06-23-redpanda-bridge-queries-in-redpanda-sql).

Relationship to small file problem

The small file problem directly undermines predicate pushdown: thousands of tiny files not only multiply S3 request costs but also defeat statistics-based pruning. Flushing less often (enabled by flush/freshness decoupling) produces the large files that make pushdown effective.

Probabilistic pushdown vs. a definitive index

Predicate pushdown (row-group min/max, Bloom filters, PageIndex) is probabilistic: it narrows what must be scanned but still requires a scan of the survivors. For selective point queries by key, Spotify's RAP instead uses a definitive external index that returns the exact files and rows, eliminating the scan entirely. The two compose — Bloom filters narrow the candidate file set, then the external index resolves exact locations — but they answer different questions: pushdown skips, the index locates (Source: sources/2026-07-27-spotify-indexing-the-data-lake-for-online-point-queries-e2d11367).

Seen in

Merged aliases

  • partition-pruning
Last updated · 766 distilled / 2,225 read