IN LIST: add Float16 bitmap filter - #23311
Merged
alamb merged 1 commit intoJul 7, 2026
Merged
Conversation
This was referenced Jul 3, 2026
geoffreyclaude
force-pushed
the
perf/in_list_float16_bitmap_filter
branch
from
July 3, 2026 18:17
2630cbd to
6f7992e
Compare
Contributor
Author
|
@alamb I forgot to mention |
alamb
approved these changes
Jul 7, 2026
alamb
left a comment
Contributor
There was a problem hiding this comment.
makes sense to me -- thanks @geoffreyclaude
rohankumardubey
pushed a commit
to rohankumardubey/datafusion
that referenced
this pull request
Jul 7, 2026
## Which issue does this PR close? - part of apache#23307 ## Rationale for this change As @geoffreyclaude adds additional special implementations for IN lists we should keep up with slt test coverage. Recently we added coverage for floats - apache#23311 ## What changes are included in this PR? This adds sqllogictest coverage for Float16, Float32, and Float64 `IN` list predicates in `in_list.slt`, mirroring the existing integer `IN` list coverage shape: ## Are these changes tested? Yes -- all tests ## Are there any user-facing changes? No. This only adds test coverage.
pull Bot
pushed a commit
to TCeason/arrow-datafusion
that referenced
this pull request
Jul 26, 2026
## Which issue does this PR close? - Part of apache#19241. - Stacked on [apache#23311](apache#23311). - Next in stack: apache#23015. - Extracted from apache#19390. ## Rationale for this change For very small `IN` lists, building or probing a hash table can be more work than just comparing the input value with each constant. For example, for `x IN (10, 20, 30)`, the fast path can behave like: ```text x == 10 OR x == 20 OR x == 30 ``` Because the list is tiny, those comparisons are cheap. The implementation stores the constants in a fixed-size array and checks them with a compact comparison chain. “Branchless” here means the comparisons are combined without stopping at the first match. That can be faster for these small fixed-width lists because the CPU gets a predictable sequence of simple operations instead of hash-table setup and probe logic. For primitive values that are not already plain unsigned integers, this PR keeps the logical Arrow type explicit and uses a matching same-width comparison representation only inside the branchless filter. For example, `Float16` uses `UInt16` storage, `Float32` uses `UInt32` storage, and `TimestampNanosecond` uses `UInt64` storage. `Decimal128` and `IntervalMonthDayNano` use their own 16-byte native representation. This preserves bit-pattern equality while relying on Arrow's native primitive compatibility rules: timestamp timezone metadata and Decimal128 precision/scale metadata may differ, while incompatible primitive representations remain rejected. ## What changes are included in this PR? - Adds a const-generic `BranchlessFilter` for small primitive `IN` lists. - Adds thresholds for when this path is used: - up to 16 values for 1-byte types - up to 8 values for 2-byte types - up to 32 values for 4-byte types - up to 16 values for 8-byte types - up to 4 values for 16-byte types - Keeps dispatch concrete and explicit in `strategy.rs`. - Maps each optimized logical type to the comparison representation used by the branchless filter: - `Int8` -> `UInt8` - `Int16`, `Float16` -> `UInt16` - `Int32`, `Float32`, `Date32`, `Time32` -> `UInt32` - `Int64`, `Float64`, `Date64`, `Time64`, `Timestamp`, `Duration` -> `UInt64` - `Decimal128`, `IntervalMonthDayNano` -> their native 16-byte representation - Leaves larger 1-byte and 2-byte lists on the existing bitmap filters. - Leaves larger 4-byte and 8-byte lists on the existing hash/generic paths. - Leaves wider primitive types such as `Decimal256` and unsupported complex types on the generic path. - Keeps the same `IN` / `NOT IN` null behavior as the rest of the stack. - Adds focused coverage for branchless null handling, signed boundary values, slices, Float16/Float32/Float64 bit patterns, compatible timestamp/Decimal128 metadata, incompatible timestamp units, IntervalMonthDayNano values, and same-width wrong-type probe rejection. ## Are these changes tested? Yes. - `cargo fmt --all -- --check` - `cargo test -p datafusion-physical-expr expressions::in_list --lib` - `cargo test -p datafusion-physical-expr --bench in_list_strategy --no-run` - `cargo clippy --all-targets --all-features -- -D warnings` ## Are there any user-facing changes? No. This is an internal performance optimization only. ## Local benchmark snapshot Built and run with `release-nonlto`, filtered to the relevant small primitive-list rows: ```bash cargo bench -p datafusion-physical-expr --profile release-nonlto --bench in_list_strategy -- <filter> --save-baseline <baseline> ``` Filters used: `narrow_integer`, `primitive/i32/small_list`, `primitive/i64/small_list`, `f32/small_list`, `timestamp_ns/small_list`, and `interval_month_day_nano/small_list`. Method: directly compared Criterion's raw sample minima (`min(time / iterations)`) from `sample.json`. Lower is better; changes within +/-5% are treated as noise. Compared baselines: [apache#23311](apache#23311) -> [apache#23014](apache#23014) Relevant scope: small primitive-list rows. Summary: 39 relevant rows, 28 faster, 0 slower, 11 within +/-5%. Largest relevant deltas: | Benchmark | Before | After | Change | |---|---:|---:|---:| | `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us | -93.2% (14.69x faster) | | `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0% (11.15x faster) | | `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us | -90.5% (10.58x faster) | | `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us | -90.5% (10.54x faster) | | `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us | -83.8% (6.16x faster) | | `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x faster) | | `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us | -82.1% (5.59x faster) | | `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us | -81.2% (5.31x faster) | | `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26 us | -77.3% (4.41x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35 us | 7.32 us | -75.1% (4.01x faster) | | `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us | -74.0% (3.84x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89 us | 7.31 us | -71.8% (3.54x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` | 26.05 us | 7.42 us | -71.5% (3.51x faster) | | `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us | 15.52 us | -70.7% (3.41x faster) | | `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8% (2.92x faster) | | `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us | -60.1% (2.50x faster) | <details> <summary>Full relevant table (39 rows)</summary> | Benchmark | Before | After | Change | |---|---:|---:|---:| | `narrow_integer/u8/list=4/match=0%` | 3.86 us | 2.79 us | -27.8% (1.38x faster) | | `narrow_integer/u8/list=4/match=50%` | 3.84 us | 2.78 us | -27.7% (1.38x faster) | | `narrow_integer/u8/list=16/match=0%` | 3.88 us | 3.85 us | -0.8% (within noise) | | `narrow_integer/u8/list=16/match=50%` | 3.84 us | 3.86 us | +0.5% (within noise) | | `narrow_integer/i16/list=4/match=0%` | 3.93 us | 3.18 us | -19.1% (1.24x faster) | | `narrow_integer/i16/list=4/match=50%` | 3.92 us | 3.16 us | -19.5% (1.24x faster) | | `narrow_integer/i16/list=64/match=0%` | 3.96 us | 3.82 us | -3.5% (within noise) | | `narrow_integer/i16/list=64/match=50%` | 3.91 us | 3.80 us | -2.9% (within noise) | | `narrow_integer/i16/list=256/match=0%` | 3.90 us | 3.81 us | -2.5% (within noise) | | `narrow_integer/i16/list=256/match=50%` | 3.97 us | 3.81 us | -4.1% (within noise) | | `narrow_integer/f16/list=4/match=0%` | 3.87 us | 3.16 us | -18.5% (1.23x faster) | | `narrow_integer/f16/list=4/match=50%` | 3.94 us | 3.15 us | -20.2% (1.25x faster) | | `narrow_integer/f16/list=64/match=0%` | 3.87 us | 3.84 us | -0.6% (within noise) | | `narrow_integer/f16/list=64/match=50%` | 3.93 us | 3.85 us | -1.9% (within noise) | | `narrow_integer/f16/list=256/match=0%` | 3.90 us | 3.84 us | -1.5% (within noise) | | `narrow_integer/f16/list=256/match=50%` | 3.87 us | 3.91 us | +1.2% (within noise) | | `nulls/narrow_integer/u8/list=16/match=50%/nulls=20%` | 3.92 us | 4.02 us | +2.5% (within noise) | | `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us | -82.1% (5.59x faster) | | `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us | -90.5% (10.58x faster) | | `primitive/i32/small_list/list=32/match=0%` | 16.34 us | 13.33 us | -18.5% (1.23x faster) | | `primitive/i32/small_list/list=32/match=50%` | 31.17 us | 13.31 us | -57.3% (2.34x faster) | | `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26 us | -77.3% (4.41x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35 us | 7.32 us | -75.1% (4.01x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` | 26.05 us | 7.42 us | -71.5% (3.51x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89 us | 7.31 us | -71.8% (3.54x faster) | | `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us | -81.2% (5.31x faster) | | `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us | -90.5% (10.54x faster) | | `primitive/i64/small_list/list=16/match=0%` | 16.34 us | 11.93 us | -27.0% (1.37x faster) | | `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us | -60.1% (2.50x faster) | | `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x faster) | | `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0% (11.15x faster) | | `f32/small_list/list=32/match=0%` | 22.05 us | 13.43 us | -39.1% (1.64x faster) | | `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8% (2.92x faster) | | `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us | -83.8% (6.16x faster) | | `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us | -93.2% (14.69x faster) | | `timestamp_ns/small_list/list=16/match=0%` | 19.73 us | 12.07 us | -38.8% (1.63x faster) | | `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us | -74.0% (3.84x faster) | | `interval_month_day_nano/small_list/list=4/match=0%` | 20.12 us | 13.20 us | -34.4% (1.52x faster) | | `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us | 15.52 us | -70.7% (3.41x faster) | </details> --------- Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
Rationale for this change
#23299 extends the bitmap
INfilter to the signed 1-byte and 2-byte integer types by handling each logical Arrow type directly.Float16is the remaining 2-byte primitive type that can use the same compact bitmap idea: it has 65,536 possible bit patterns, so an 8 KiB bitmap can represent every possible value.This PR follows the same direct typed shape as #23299. It does not reinterpret whole arrays as
UInt16; instead, theFloat16bitmap filter maps each value to its IEEE-754 half-precision bit pattern withto_bits(). That keeps the logical array type intact while preserving bit-pattern equality semantics, including distinct NaN payloads and+0.0versus-0.0.What changes are included in this PR?
Float16Typesupport to the existingBitmapFilter.DataType::Float16constant-list filtering to that bitmap path.Float16.NOT IN,+0.0/-0.0, and NaN payload bit patterns.in_list_strategybenchmark rows forFloat16.Are these changes tested?
Yes.
cargo fmt --allcargo test -p datafusion-physical-expr bitmap_filter_f16 --libcargo test -p datafusion-physical-expr test_in_list_from_array_type_combinations --libcargo test -p datafusion-physical-expr --bench in_list_strategy --no-runcargo clippy --all-targets --all-features -- -D warningsAre there any user-facing changes?
No. This is an internal performance optimization only.
Local benchmark snapshot
Built and run with
release-nonlto, filtered to the new Float16 rows:Compared baselines: #23299 -> #23311
Method: directly compared Criterion's raw sample minima (
min(time / iterations)) fromsample.json. Lower is better; changes within +/-5% are treated as noise.Summary: 6 relevant rows, 6 faster, 0 slower, 0 within +/-5%.
narrow_integer/f16/list=4/match=0%narrow_integer/f16/list=4/match=50%narrow_integer/f16/list=64/match=0%narrow_integer/f16/list=64/match=50%narrow_integer/f16/list=256/match=0%narrow_integer/f16/list=256/match=50%