CONCEPT Cited by 3 sources
Predicate pushdown¶
Predicate pushdown is the query-optimization technique of pushing filter predicates (WHERE clauses) down into the storage layer so that irrelevant data is never read into the execution engine. In columnar formats like Parquet, this means consulting per-row-group min/max statistics to skip entire row groups that cannot satisfy the predicate.
Effectiveness depends on file size¶
Row group statistics work best when row groups are substantial. Tiny Parquet files (e.g. 500 KB written every 30 seconds from a streaming flush) often have row groups too small to prune meaningfully — the statistics span so little data that most predicates match anyway. Larger files (32+ MB) with well-populated row groups enable significant data-skipping savings (Source: sources/2026-06-23-redpanda-bridge-queries-in-redpanda-sql).
Relationship to small file problem¶
The small file problem directly undermines predicate pushdown: thousands of tiny files not only multiply S3 request costs but also defeat statistics-based pruning. Flushing less often (enabled by flush/freshness decoupling) produces the large files that make pushdown effective.
Probabilistic pushdown vs. a definitive index¶
Predicate pushdown (row-group min/max, Bloom filters, PageIndex) is probabilistic: it narrows what must be scanned but still requires a scan of the survivors. For selective point queries by key, Spotify's RAP instead uses a definitive external index that returns the exact files and rows, eliminating the scan entirely. The two compose — Bloom filters narrow the candidate file set, then the external index resolves exact locations — but they answer different questions: pushdown skips, the index locates (Source: sources/2026-07-27-spotify-indexing-the-data-lake-for-online-point-queries-e2d11367).
Seen in¶
- systems/apache-iceberg — Iceberg manifest-level statistics enable partition pruning and data file skipping
- systems/apache-parquet — row group footer statistics power column-level pushdown
- systems/redpanda-sql — benefits from large Parquet files written by decoupled flush intervals
- systems/random-access-parquet — contrasts probabilistic pushdown/Bloom filters with a definitive external index for point queries; hoisted values enable index-level predicate pushdown before any storage read (Source: sources/2026-07-27-spotify-indexing-the-data-lake-for-online-point-queries-e2d11367)
- sources/2026-06-23-redpanda-bridge-queries-in-redpanda-sql — describes effectiveness scaling with file size
- sources/2026-07-29-redpanda-single-query-scaling-in-redpanda-sql — demonstrates constant ~4 s latency from 100 GB to 1 TB when time-range filter prunes via Iceberg manifest min/max (reads 76 of 4,961 files at 1 TB)
Merged aliases¶
partition-pruning