perf[gpu]: generate FixedSizeList offsets on device - #8458
Merged
CodSpeed HQ / CodSpeed Performance Analysis
succeeded
Jun 17, 2026 in 0s
Performance Gate Passed
⚠️ Unknown Walltime execution environment detected
Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.
For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.
⚡ 9 improved benchmarks
✅ 1536 untouched benchmarks
⏩ 10 skipped benchmarks1
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | Simulation | chunked_bool_canonical_into[(1000, 10)] |
35.6 µs | 20.6 µs | +72.88% |
| ⚡ | Simulation | baseline_eq[4, 65536] |
243.5 µs | 185.7 µs | +31.08% |
| ⚡ | Simulation | baseline_eq[16, 65536] |
288.4 µs | 230.7 µs | +25.04% |
| ⚡ | Simulation | baseline_lt[16, 65536] |
303.6 µs | 246.1 µs | +23.38% |
| ⚡ | Simulation | chunked_varbinview_canonical_into[(1000, 10)] |
198.2 µs | 162 µs | +22.38% |
| ⚡ | Simulation | chunked_varbinview_into_canonical[(1000, 10)] |
213.2 µs | 177 µs | +20.45% |
| ⚡ | Simulation | bitwise_not_vortex_buffer_mut[128] |
244.4 ns | 215.3 ns | +13.55% |
| ⚡ | Simulation | encode_varbin[(1000, 2)] |
176.4 µs | 158.3 µs | +11.42% |
| ⚡ | Simulation | bitwise_not_vortex_buffer_mut[1024] |
304.7 ns | 275.6 ns | +10.58% |
Tip
Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ad/list-offset-device (1db1af1) with develop (4652f29)
Footnotes
-
10 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
Loading