Skip to content

IN LIST: add Float16 bitmap filter - #23311

Merged
alamb merged 1 commit into
apache:mainfrom
geoffreyclaude:perf/in_list_float16_bitmap_filter
Jul 7, 2026
Merged

IN LIST: add Float16 bitmap filter#23311
alamb merged 1 commit into
apache:mainfrom
geoffreyclaude:perf/in_list_float16_bitmap_filter

Conversation

@geoffreyclaude

@geoffreyclaude geoffreyclaude commented Jul 3, 2026

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

#23299 extends the bitmap IN filter to the signed 1-byte and 2-byte integer types by handling each logical Arrow type directly. Float16 is the remaining 2-byte primitive type that can use the same compact bitmap idea: it has 65,536 possible bit patterns, so an 8 KiB bitmap can represent every possible value.

This PR follows the same direct typed shape as #23299. It does not reinterpret whole arrays as UInt16; instead, the Float16 bitmap filter maps each value to its IEEE-754 half-precision bit pattern with to_bits(). That keeps the logical array type intact while preserving bit-pattern equality semantics, including distinct NaN payloads and +0.0 versus -0.0.

What changes are included in this PR?

  • Adds Float16Type support to the existing BitmapFilter.
  • Routes DataType::Float16 constant-list filtering to that bitmap path.
  • Extends the existing type-combination coverage to include Float16.
  • Adds focused coverage for slices, nulls, NOT IN, +0.0 / -0.0, and NaN payload bit patterns.
  • Adds focused in_list_strategy benchmark rows for Float16.

Are these changes tested?

Yes.

  • cargo fmt --all
  • cargo test -p datafusion-physical-expr bitmap_filter_f16 --lib
  • cargo test -p datafusion-physical-expr test_in_list_from_array_type_combinations --lib
  • cargo test -p datafusion-physical-expr --bench in_list_strategy --no-run
  • cargo clippy --all-targets --all-features -- -D warnings

Are there any user-facing changes?

No. This is an internal performance optimization only.

Local benchmark snapshot

Built and run with release-nonlto, filtered to the new Float16 rows:

cargo bench -p datafusion-physical-expr --profile release-nonlto --bench in_list_strategy -- narrow_integer/f16 --save-baseline <baseline>

Compared baselines: #23299 -> #23311

Method: directly compared Criterion's raw sample minima (min(time / iterations)) from sample.json. Lower is better; changes within +/-5% are treated as noise.

Summary: 6 relevant rows, 6 faster, 0 slower, 0 within +/-5%.

Benchmark #23299 #23311 Change
narrow_integer/f16/list=4/match=0% 19.582 us 3.911 us -80.0% (5.01x faster)
narrow_integer/f16/list=4/match=50% 44.138 us 3.871 us -91.2% (11.40x faster)
narrow_integer/f16/list=64/match=0% 19.977 us 3.878 us -80.6% (5.15x faster)
narrow_integer/f16/list=64/match=50% 55.792 us 3.903 us -93.0% (14.29x faster)
narrow_integer/f16/list=256/match=0% 21.727 us 3.885 us -82.1% (5.59x faster)
narrow_integer/f16/list=256/match=50% 51.737 us 3.918 us -92.4% (13.21x faster)

@geoffreyclaude
geoffreyclaude force-pushed the perf/in_list_float16_bitmap_filter branch from 2630cbd to 6f7992e Compare July 3, 2026 18:17
@geoffreyclaude

geoffreyclaude commented Jul 3, 2026

Copy link
Copy Markdown
Contributor Author

@alamb I forgot to mention Float16 in your Optimize Int8 and Int16 integer IN filters . This should cover the gap.

@alamb alamb left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

makes sense to me -- thanks @geoffreyclaude

@alamb
alamb added this pull request to the merge queue Jul 7, 2026
Merged via the queue into apache:main with commit 34274f3 Jul 7, 2026
39 checks passed
rohankumardubey pushed a commit to rohankumardubey/datafusion that referenced this pull request Jul 7, 2026
## Which issue does this PR close?

- part of apache#23307

## Rationale for this change

As @geoffreyclaude adds additional special implementations for IN lists
we should keep up with slt test coverage. Recently we added coverage for
floats
- apache#23311

## What changes are included in this PR?

This adds sqllogictest coverage for Float16, Float32, and Float64 `IN`
list predicates in `in_list.slt`, mirroring the existing integer `IN`
list coverage shape:


## Are these changes tested?
Yes -- all tests
## Are there any user-facing changes?

No. This only adds test coverage.
pull Bot pushed a commit to TCeason/arrow-datafusion that referenced this pull request Jul 26, 2026
## Which issue does this PR close?

- Part of apache#19241.
- Stacked on [apache#23311](apache#23311).
- Next in stack: apache#23015.
- Extracted from apache#19390.

## Rationale for this change

For very small `IN` lists, building or probing a hash table can be more
work than just comparing the input value with each constant.

For example, for `x IN (10, 20, 30)`, the fast path can behave like:

```text
x == 10 OR x == 20 OR x == 30
```

Because the list is tiny, those comparisons are cheap. The
implementation stores the constants in a fixed-size array and checks
them with a compact comparison chain.

“Branchless” here means the comparisons are combined without stopping at
the first match. That can be faster for these small fixed-width lists
because the CPU gets a predictable sequence of simple operations instead
of hash-table setup and probe logic.

For primitive values that are not already plain unsigned integers, this
PR keeps the logical Arrow type explicit and uses a matching same-width
comparison representation only inside the branchless filter. For
example, `Float16` uses `UInt16` storage, `Float32` uses `UInt32`
storage, and `TimestampNanosecond` uses `UInt64` storage. `Decimal128`
and `IntervalMonthDayNano` use their own 16-byte native representation.
This preserves bit-pattern equality while relying on Arrow's native
primitive compatibility rules: timestamp timezone metadata and
Decimal128 precision/scale metadata may differ, while incompatible
primitive representations remain rejected.

## What changes are included in this PR?

- Adds a const-generic `BranchlessFilter` for small primitive `IN`
lists.
- Adds thresholds for when this path is used:
  - up to 16 values for 1-byte types
  - up to 8 values for 2-byte types
  - up to 32 values for 4-byte types
  - up to 16 values for 8-byte types
  - up to 4 values for 16-byte types
- Keeps dispatch concrete and explicit in `strategy.rs`.
- Maps each optimized logical type to the comparison representation used
by the branchless filter:
  - `Int8` -> `UInt8`
  - `Int16`, `Float16` -> `UInt16`
  - `Int32`, `Float32`, `Date32`, `Time32` -> `UInt32`
- `Int64`, `Float64`, `Date64`, `Time64`, `Timestamp`, `Duration` ->
`UInt64`
- `Decimal128`, `IntervalMonthDayNano` -> their native 16-byte
representation
- Leaves larger 1-byte and 2-byte lists on the existing bitmap filters.
- Leaves larger 4-byte and 8-byte lists on the existing hash/generic
paths.
- Leaves wider primitive types such as `Decimal256` and unsupported
complex types on the generic path.
- Keeps the same `IN` / `NOT IN` null behavior as the rest of the stack.
- Adds focused coverage for branchless null handling, signed boundary
values, slices, Float16/Float32/Float64 bit patterns, compatible
timestamp/Decimal128 metadata, incompatible timestamp units,
IntervalMonthDayNano values, and same-width wrong-type probe rejection.

## Are these changes tested?

Yes.

- `cargo fmt --all -- --check`
- `cargo test -p datafusion-physical-expr expressions::in_list --lib`
- `cargo test -p datafusion-physical-expr --bench in_list_strategy
--no-run`
- `cargo clippy --all-targets --all-features -- -D warnings`

## Are there any user-facing changes?

No. This is an internal performance optimization only.

## Local benchmark snapshot

Built and run with `release-nonlto`, filtered to the relevant small
primitive-list rows:

```bash
cargo bench -p datafusion-physical-expr --profile release-nonlto --bench in_list_strategy -- <filter> --save-baseline <baseline>
```

Filters used: `narrow_integer`, `primitive/i32/small_list`,
`primitive/i64/small_list`, `f32/small_list`, `timestamp_ns/small_list`,
and `interval_month_day_nano/small_list`.

Method: directly compared Criterion's raw sample minima (`min(time /
iterations)`) from `sample.json`. Lower is better; changes within +/-5%
are treated as noise.

Compared baselines:
[apache#23311](apache#23311) ->
[apache#23014](apache#23014)

Relevant scope: small primitive-list rows.

Summary: 39 relevant rows, 28 faster, 0 slower, 11 within +/-5%.

Largest relevant deltas:

| Benchmark | Before | After | Change |
|---|---:|---:|---:|
| `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us |
-93.2% (14.69x faster) |
| `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0%
(11.15x faster) |
| `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us |
-90.5% (10.58x faster) |
| `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us |
-90.5% (10.54x faster) |
| `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us |
-83.8% (6.16x faster) |
| `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x
faster) |
| `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us |
-82.1% (5.59x faster) |
| `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us |
-81.2% (5.31x faster) |
| `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26
us | -77.3% (4.41x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35
us | 7.32 us | -75.1% (4.01x faster) |
| `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us |
-74.0% (3.84x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89
us | 7.31 us | -71.8% (3.54x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` |
26.05 us | 7.42 us | -71.5% (3.51x faster) |
| `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us |
15.52 us | -70.7% (3.41x faster) |
| `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8%
(2.92x faster) |
| `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us |
-60.1% (2.50x faster) |

<details>
<summary>Full relevant table (39 rows)</summary>

| Benchmark | Before | After | Change |
|---|---:|---:|---:|
| `narrow_integer/u8/list=4/match=0%` | 3.86 us | 2.79 us | -27.8%
(1.38x faster) |
| `narrow_integer/u8/list=4/match=50%` | 3.84 us | 2.78 us | -27.7%
(1.38x faster) |
| `narrow_integer/u8/list=16/match=0%` | 3.88 us | 3.85 us | -0.8%
(within noise) |
| `narrow_integer/u8/list=16/match=50%` | 3.84 us | 3.86 us | +0.5%
(within noise) |
| `narrow_integer/i16/list=4/match=0%` | 3.93 us | 3.18 us | -19.1%
(1.24x faster) |
| `narrow_integer/i16/list=4/match=50%` | 3.92 us | 3.16 us | -19.5%
(1.24x faster) |
| `narrow_integer/i16/list=64/match=0%` | 3.96 us | 3.82 us | -3.5%
(within noise) |
| `narrow_integer/i16/list=64/match=50%` | 3.91 us | 3.80 us | -2.9%
(within noise) |
| `narrow_integer/i16/list=256/match=0%` | 3.90 us | 3.81 us | -2.5%
(within noise) |
| `narrow_integer/i16/list=256/match=50%` | 3.97 us | 3.81 us | -4.1%
(within noise) |
| `narrow_integer/f16/list=4/match=0%` | 3.87 us | 3.16 us | -18.5%
(1.23x faster) |
| `narrow_integer/f16/list=4/match=50%` | 3.94 us | 3.15 us | -20.2%
(1.25x faster) |
| `narrow_integer/f16/list=64/match=0%` | 3.87 us | 3.84 us | -0.6%
(within noise) |
| `narrow_integer/f16/list=64/match=50%` | 3.93 us | 3.85 us | -1.9%
(within noise) |
| `narrow_integer/f16/list=256/match=0%` | 3.90 us | 3.84 us | -1.5%
(within noise) |
| `narrow_integer/f16/list=256/match=50%` | 3.87 us | 3.91 us | +1.2%
(within noise) |
| `nulls/narrow_integer/u8/list=16/match=50%/nulls=20%` | 3.92 us | 4.02
us | +2.5% (within noise) |
| `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us |
-82.1% (5.59x faster) |
| `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us |
-90.5% (10.58x faster) |
| `primitive/i32/small_list/list=32/match=0%` | 16.34 us | 13.33 us |
-18.5% (1.23x faster) |
| `primitive/i32/small_list/list=32/match=50%` | 31.17 us | 13.31 us |
-57.3% (2.34x faster) |
| `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26
us | -77.3% (4.41x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35
us | 7.32 us | -75.1% (4.01x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` |
26.05 us | 7.42 us | -71.5% (3.51x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89
us | 7.31 us | -71.8% (3.54x faster) |
| `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us |
-81.2% (5.31x faster) |
| `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us |
-90.5% (10.54x faster) |
| `primitive/i64/small_list/list=16/match=0%` | 16.34 us | 11.93 us |
-27.0% (1.37x faster) |
| `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us |
-60.1% (2.50x faster) |
| `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x
faster) |
| `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0%
(11.15x faster) |
| `f32/small_list/list=32/match=0%` | 22.05 us | 13.43 us | -39.1%
(1.64x faster) |
| `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8%
(2.92x faster) |
| `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us |
-83.8% (6.16x faster) |
| `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us |
-93.2% (14.69x faster) |
| `timestamp_ns/small_list/list=16/match=0%` | 19.73 us | 12.07 us |
-38.8% (1.63x faster) |
| `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us |
-74.0% (3.84x faster) |
| `interval_month_day_nano/small_list/list=4/match=0%` | 20.12 us |
13.20 us | -34.4% (1.52x faster) |
| `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us |
15.52 us | -70.7% (3.41x faster) |

</details>

---------

Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

physical-expr Changes to the physical-expr crates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants