Skip to content

perf(dc): pack the bit-backend band sign planes once per track! call#206

Merged
zsoerenm merged 2 commits into
masterfrom
perf/pack-band-once
Jul 22, 2026
Merged

perf(dc): pack the bit-backend band sign planes once per track! call#206
zsoerenm merged 2 commits into
masterfrom
perf/pack-band-once

Conversation

@zsoerenm

@zsoerenm zsoerenm commented Jul 22, 2026

Copy link
Copy Markdown
Member

Summary

The one-/two-bit backends pack the whole measurement buffer's shared band sign planes (mrband/miband, plus the magnitude planes for two-bit) into the correlator's preallocated BandBuffers on every downconvert_and_correlate! call when a group holds more than one satellite. track! calls it once per integration step, so packing was O(steps × buffer) for an O(buffer) job.

What changed

New samples_unchanged::Bool = false kwarg on downconvert_and_correlate(!): a promise that every band's sample buffer holds the same content as on the previous call with this dc, allowing backends to reuse sample-derived caches. track! passes it on every loop iteration after the first, so the shared band pack now runs exactly once per call. The float backends ignore the flag; direct callers keep the safe default (false ⇒ repack), so real-time loops that refill the same buffer each call are unaffected.

Benchmark

The gain scales with how many integration steps one track! call spans: nil for 1 ms real-time-style buffers (one step ⇒ one pack either way — which is why the repo's benchmark suite, whose multi-sat cases feed 1 ms buffers, never flagged this), growing with buffer length for post-processing-style calls. 8-sat GPS L1 C/A Complex{Int16} capture at 5 MHz, single-threaded backends, best of 20 full track! calls:

buffer one-bit two-bit
1 ms (5 k samples) 6.1 → 6.3 µs (no change) 7.5 → 7.5 µs (no change)
20 ms (100 k) 151 → 107 µs (~1.4×) 207 → 142 µs (~1.5×)
100 ms (500 k) 2108 → 559 µs (~3.8×) 2849 → 745 µs (~3.8×)

Testing

  • New benchmark scenario "GPS L1CA, 8 sats @ 5 MHz, 100 ms buffer" in the existing track! Int16 vs Float32 group, so this PR's benchmark CI table shows the before/after directly (the 1 ms scenarios are unaffected by design).
  • downconvert_and_correlate_onebit / _twobit / _int16 / downconvert_and_correlate, track, track_in_place (0-byte steady-state assertions), and multi_signal suites pass locally.
  • JuliaFormatter 2.8.5 clean.

Note: #205 contains the equivalent fix adapted to its chunked two-pass loop (where the default 1 ms chunking makes the redundant packing bite even for long buffers at the default settings); whichever merges second needs a trivial rebase of these hunks.

🤖 Generated with Claude Code

The one-/two-bit backends re-packed the WHOLE measurement buffer's
shared band sign planes (>1 sat) on every downconvert_and_correlate!
call — i.e. once per integration step inside track!'s loop — making
packing O(steps x buffer) instead of O(buffer). New samples_unchanged
kwarg (default false) promises the band buffers hold the same content
as on the previous call with this dc, letting backends reuse
sample-derived caches; track! passes it on every loop iteration after
the first, so the pack now happens exactly once per call. The float
backends ignore it, and direct downconvert_and_correlate(!) callers
keep the safe default.

The gain scales with how many integration steps one track! call spans:
nil for the 1 ms real-time-style buffers the benchmark suite uses (one
step = one pack either way), and growing with buffer length for
post-processing-style calls. 8-sat GPS L1CA Complex{Int16} capture at
5 MHz, single-threaded backends, best of 20 track! calls:

    buffer          one-bit             two-bit
    1 ms (5k)       6.1 -> 6.3 us       7.5 -> 7.5 us
    20 ms (100k)    151 -> 107 us       207 -> 142 us
    100 ms (500k)   2108 -> 559 us      2849 -> 745 us

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@zsoerenm
zsoerenm force-pushed the perf/pack-band-once branch 2 times, most recently from a3e41cd to 2991178 Compare July 22, 2026 12:33
@github-actions

github-actions Bot commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Benchmark Results (minimum time) — macos-14

Reporting the minimum over all samples (robust to shared-runner contention), not the median.

Alternative backends vs Float32 (track!, PR head)

Legend — backends: F32 Float32 (default) · I16 Int16 · 1b OneBit · 2b TwoBit (2-bit measurement + 2-bit carrier). Time columns are the minimum track! time; ×B = F32 / B (so >1 ⇒ backend B is faster than Float32), ✅ ≥ 5 % faster, ⚠️ ≥ 5 % slower. 1b/2b are BPSK-only, so their cells are blank for CBOC (Galileo E1B) scenarios.

Scenario F32 I16 1b 2b ×I16 ×1b ×2b
4-antenna @ 5 MHz 12.3 μs 6.27 μs 4.02 μs 8.75 μs 1.96 ✅ 3.06 ✅ 1.4 ✅
GPS L1CA, 8 sats @ 40 MHz 368.0 μs 118.0 μs 58.1 μs 107.0 μs 3.12 ✅ 6.34 ✅ 3.44 ✅
GPS L1CA, 8 sats @ 5 MHz, 100 ms buffer 5.2 ms 1.94 ms 1.21 ms 1.95 ms 2.67 ✅ 4.29 ✅ 2.66 ✅
GPS L1CA, 8 sats @ 5 MHz 53.1 μs 22.1 μs 15.8 μs 23.4 μs 2.4 ✅ 3.36 ✅ 2.27 ✅
Galileo E1B, 4 sats @ 25 MHz 141.0 μs 59.2 μs 2.39 ✅
dynamic taps @ 5 MHz (kernel) 6.68 μs 2.23 μs 1.48 μs 2.47 μs 3.0 ✅ 4.51 ✅ 2.71 ✅
multi-signal N=3 @ 5 MHz 10.6 μs 5.08 μs 3.69 μs 5.94 μs 2.08 ✅ 2.87 ✅ 1.78 ✅
Time benchmarks (base vs PR head)

Ratio = 5f03cd2… / a0be4ef…: >1 means the PR is faster. ✅ ≥ 5 % faster, ⚠️ ≥ 5 % slower. A blank cell means the benchmark exists on only one revision (🆕 = new on the PR, 🗑 = removed).

5f03cd2 a0be4ef 5f03cd2… / a0be4ef
downconvert and correlate/CPU/Float32 2.62 μs 2.57 μs 1.02
downconvert and correlate/CPU/Float32 4ant 5.08 μs 5.05 μs 1.01
downconvert and correlate/CPU/Float64 2.81 μs 2.81 μs 1.0
downconvert and correlate/CPU/Int16 2.8 μs 2.73 μs 1.03
downconvert and correlate/CPU/Int16 4ant 5.01 μs 5.0 μs 1.0
downconvert and correlate/CPU/Int32 2.75 μs 2.57 μs 1.07 ✅
fused kernel/1-ant dynamic taps 2.62 μs 2.56 μs 1.02
fused kernel/1-ant static taps 2.06 μs 2.06 μs 1.0
fused kernel/4-ant dynamic taps 8.32 μs 7.79 μs 1.07 ✅
fused kernel/4-ant static taps 4.6 μs 4.44 μs 1.04
fused tuple kernel/1-ant N=2 3.28 μs 3.28 μs 1.0
fused tuple kernel/1-ant N=3 3.41 μs 3.42 μs 0.998
fused tuple kernel/2-ant N=2 6.37 μs 6.31 μs 1.01
fused tuple kernel/2-ant N=3 6.02 μs 5.62 μs 1.07 ✅
fused tuple kernel/4-ant N=2 12.4 μs 11.6 μs 1.07 ✅
fused tuple kernel/4-ant N=3 10.8 μs 10.5 μs 1.03
track/1. Float32/2K – track 2.8 μs 2.91 μs 0.964
track/1. Float32/2K – track! 2.88 μs 2.86 μs 1.01
track/2. L1 8sat/5K – track 51.8 μs 53.2 μs 0.974
track/2. L1 8sat/5K – track! 51.3 μs 49.6 μs 1.03
track/2. L1 8sat/5K – track! Int16 21.5 μs 22.7 μs 0.947 ⚠️
track/2. L1 8sat/5K – track! OneBit 15.8 μs 16.6 μs 0.952
track/2. L1 8sat/5K – track!-threaded 52.0 μs 51.5 μs 1.01
track/2. L1 8sat/5K – track-threaded 51.8 μs 53.0 μs 0.978
track/3. E1B 4sat/25K – track 142.0 μs 140.0 μs 1.01
track/3. E1B 4sat/25K – track! 142.0 μs 140.0 μs 1.01
track/3. E1B 4sat/25K – track! Int16 60.9 μs 58.5 μs 1.04
track/3. E1B 4sat/25K – track!-threaded 140.0 μs 136.0 μs 1.04
track/3. E1B 4sat/25K – track-threaded 143.0 μs 141.0 μs 1.01
track/4. 8L1+8E1B/25K – track 508.0 μs 508.0 μs 1.0
track/4. 8L1+8E1B/25K – track! 512.0 μs 506.0 μs 1.01
track/4. 8L1+8E1B/25K – track!-threaded 514.0 μs 503.0 μs 1.02
track/4. 8L1+8E1B/25K – track-threaded 519.0 μs 507.0 μs 1.02
track/5. multi-signal N=1/5K – track 6.8 μs 6.92 μs 0.982
track/5. multi-signal N=1/5K – track! 6.72 μs 6.7 μs 1.0
track/6. multi-signal N=2/5K – track 10.8 μs 10.7 μs 1.01
track/6. multi-signal N=2/5K – track! 10.8 μs 10.9 μs 0.992
track/7. L1 8sat/500K – track! Int16 1.93 ms 1.94 ms 0.992
track/7. L1 8sat/500K – track! OneBit 5.16 ms 1.24 ms 4.16 ✅
track/7. L1 8sat/500K – track! TwoBit 9.31 ms 1.94 ms 4.8 ✅
track/7. multi-signal N=3/5K – track 11.7 μs 12.0 μs 0.972
track/7. multi-signal N=3/5K – track! 12.5 μs 12.5 μs 1.0
track/8. L1CA presync 2 blk – track! 12.6 μs 12.8 μs 0.981
track/8. L1CA presync 20 blk – track! 121.0 μs 121.0 μs 1.0
track/8. L1CA synced 2 blk – track! 12.8 μs 12.3 μs 1.04
track/8. L1CA synced 20 blk – track! 124.0 μs 121.0 μs 1.02
time_to_load 227.0 μs 333.0 μs 0.682 ⚠️
Memory benchmarks (base vs PR head)

Ratio = 5f03cd2… / a0be4ef… (bytes allocated): >1 means the PR allocates less. ✅ ≥ 5 % less, ⚠️ ≥ 5 % more. /0 mark a benchmark that drops to / picks up allocations, means both revisions allocate nothing. A blank cell means the benchmark exists on only one revision (🆕 = new on the PR, 🗑 = removed).

5f03cd2 a0be4ef 5f03cd2… / a0be4ef
downconvert and correlate/CPU/Float32 2 allocs: 576 B 2 allocs: 576 B 1.0
downconvert and correlate/CPU/Float32 4ant 2 allocs: 848 B 2 allocs: 848 B 1.0
downconvert and correlate/CPU/Float64 2 allocs: 576 B 2 allocs: 576 B 1.0
downconvert and correlate/CPU/Int16 2 allocs: 576 B 2 allocs: 576 B 1.0
downconvert and correlate/CPU/Int16 4ant 2 allocs: 848 B 2 allocs: 848 B 1.0
downconvert and correlate/CPU/Int32 2 allocs: 576 B 2 allocs: 576 B 1.0
fused kernel/1-ant dynamic taps 0 allocs: 0 B 0 allocs: 0 B
fused kernel/1-ant static taps 0 allocs: 0 B 0 allocs: 0 B
fused kernel/4-ant dynamic taps 0 allocs: 0 B 0 allocs: 0 B
fused kernel/4-ant static taps 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/1-ant N=2 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/1-ant N=3 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/2-ant N=2 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/2-ant N=3 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/4-ant N=2 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/4-ant N=3 0 allocs: 0 B 0 allocs: 0 B
track/1. Float32/2K – track 9 allocs: 944 B 9 allocs: 944 B 1.0
track/1. Float32/2K – track! 0 allocs: 0 B 0 allocs: 0 B
track/2. L1 8sat/5K – track 10 allocs: 4656 B 10 allocs: 4656 B 1.0
track/2. L1 8sat/5K – track! 0 allocs: 0 B 0 allocs: 0 B
track/2. L1 8sat/5K – track! Int16 13 allocs: 1056 B 13 allocs: 1056 B 1.0
track/2. L1 8sat/5K – track! OneBit 45 allocs: 2736 B 45 allocs: 2736 B 1.0
track/2. L1 8sat/5K – track!-threaded 0 allocs: 0 B 0 allocs: 0 B
track/2. L1 8sat/5K – track-threaded 10 allocs: 4656 B 10 allocs: 4656 B 1.0
track/3. E1B 4sat/25K – track 10 allocs: 3056 B 10 allocs: 3056 B 1.0
track/3. E1B 4sat/25K – track! 0 allocs: 0 B 0 allocs: 0 B
track/3. E1B 4sat/25K – track! Int16 5 allocs: 416 B 5 allocs: 416 B 1.0
track/3. E1B 4sat/25K – track!-threaded 0 allocs: 0 B 0 allocs: 0 B
track/3. E1B 4sat/25K – track-threaded 10 allocs: 3056 B 10 allocs: 3056 B 1.0
track/4. 8L1+8E1B/25K – track 26 allocs: 10464 B 26 allocs: 10464 B 1.0
track/4. 8L1+8E1B/25K – track! 0 allocs: 0 B 0 allocs: 0 B
track/4. 8L1+8E1B/25K – track!-threaded 0 allocs: 0 B 0 allocs: 0 B
track/4. 8L1+8E1B/25K – track-threaded 26 allocs: 10464 B 26 allocs: 10464 B 1.0
track/5. multi-signal N=1/5K – track 9 allocs: 944 B 9 allocs: 944 B 1.0
track/5. multi-signal N=1/5K – track! 0 allocs: 0 B 0 allocs: 0 B
track/6. multi-signal N=2/5K – track 9 allocs: 1408 B 9 allocs: 1408 B 1.0
track/6. multi-signal N=2/5K – track! 0 allocs: 0 B 0 allocs: 0 B
track/7. L1 8sat/500K – track! Int16 61 allocs: 28704 B 61 allocs: 28704 B 1.0
track/7. L1 8sat/500K – track! OneBit 1677 allocs: 119088 B 1677 allocs: 119088 B 1.0
track/7. L1 8sat/500K – track! TwoBit 1677 allocs: 119136 B 1677 allocs: 119136 B 1.0
track/7. multi-signal N=3/5K – track 9 allocs: 1760 B 9 allocs: 1760 B 1.0
track/7. multi-signal N=3/5K – track! 0 allocs: 0 B 0 allocs: 0 B
track/8. L1CA presync 2 blk – track! 7 allocs: 320 B 7 allocs: 320 B 1.0
track/8. L1CA presync 20 blk – track! 7 allocs: 320 B 7 allocs: 320 B 1.0
track/8. L1CA synced 2 blk – track! 7 allocs: 320 B 7 allocs: 320 B 1.0
track/8. L1CA synced 20 blk – track! 7 allocs: 320 B 7 allocs: 320 B 1.0
time_to_load 196 allocs: 13984 B 196 allocs: 13984 B 1.0

@github-actions

github-actions Bot commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Benchmark Results (minimum time) — ubuntu-latest

Reporting the minimum over all samples (robust to shared-runner contention), not the median.

Alternative backends vs Float32 (track!, PR head)

Legend — backends: F32 Float32 (default) · I16 Int16 · 1b OneBit · 2b TwoBit (2-bit measurement + 2-bit carrier). Time columns are the minimum track! time; ×B = F32 / B (so >1 ⇒ backend B is faster than Float32), ✅ ≥ 5 % faster, ⚠️ ≥ 5 % slower. 1b/2b are BPSK-only, so their cells are blank for CBOC (Galileo E1B) scenarios.

Scenario F32 I16 1b 2b ×I16 ×1b ×2b
4-antenna @ 5 MHz 11.4 μs 6.85 μs 4.75 μs 8.96 μs 1.67 ✅ 2.4 ✅ 1.27 ✅
GPS L1CA, 8 sats @ 40 MHz 337.0 μs 136.0 μs 62.2 μs 115.0 μs 2.48 ✅ 5.41 ✅ 2.93 ✅
GPS L1CA, 8 sats @ 5 MHz, 100 ms buffer 4.68 ms 2.27 ms 1.33 ms 2.08 ms 2.06 ✅ 3.5 ✅ 2.25 ✅
GPS L1CA, 8 sats @ 5 MHz 50.5 μs 27.3 μs 17.1 μs 25.3 μs 1.85 ✅ 2.95 ✅ 2.0 ✅
Galileo E1B, 4 sats @ 25 MHz 125.0 μs 58.5 μs 2.14 ✅
dynamic taps @ 5 MHz (kernel) 6.26 μs 3.51 μs 1.75 μs 2.64 μs 1.78 ✅ 3.57 ✅ 2.37 ✅
multi-signal N=3 @ 5 MHz 10.1 μs 5.51 μs 4.41 μs 6.79 μs 1.83 ✅ 2.29 ✅ 1.49 ✅
Time benchmarks (base vs PR head)

Ratio = 5f03cd2… / a0be4ef…: >1 means the PR is faster. ✅ ≥ 5 % faster, ⚠️ ≥ 5 % slower. A blank cell means the benchmark exists on only one revision (🆕 = new on the PR, 🗑 = removed).

5f03cd2 a0be4ef 5f03cd2… / a0be4ef
downconvert and correlate/CPU/Float32 2.38 μs 2.38 μs 0.999
downconvert and correlate/CPU/Float32 4ant 4.41 μs 4.41 μs 1.0
downconvert and correlate/CPU/Float64 2.76 μs 2.78 μs 0.995
downconvert and correlate/CPU/Int16 2.5 μs 2.48 μs 1.01
downconvert and correlate/CPU/Int16 4ant 5.02 μs 4.87 μs 1.03
downconvert and correlate/CPU/Int32 2.42 μs 2.4 μs 1.01
fused kernel/1-ant dynamic taps 2.29 μs 2.31 μs 0.995
fused kernel/1-ant static taps 1.91 μs 1.91 μs 0.998
fused kernel/4-ant dynamic taps 5.5 μs 5.62 μs 0.977
fused kernel/4-ant static taps 3.88 μs 3.89 μs 0.998
fused tuple kernel/1-ant N=2 2.57 μs 2.56 μs 1.0
fused tuple kernel/1-ant N=3 3.08 μs 3.13 μs 0.984
fused tuple kernel/2-ant N=2 3.92 μs 3.93 μs 0.997
fused tuple kernel/2-ant N=3 4.98 μs 4.96 μs 1.0
fused tuple kernel/4-ant N=2 6.52 μs 6.51 μs 1.0
fused tuple kernel/4-ant N=3 10.2 μs 10.2 μs 0.997
track/1. Float32/2K – track 2.6 μs 2.58 μs 1.01
track/1. Float32/2K – track! 2.64 μs 2.64 μs 1.0
track/2. L1 8sat/5K – track 48.3 μs 48.2 μs 1.0
track/2. L1 8sat/5K – track! 47.2 μs 47.5 μs 0.995
track/2. L1 8sat/5K – track! Int16 28.1 μs 26.8 μs 1.05
track/2. L1 8sat/5K – track! OneBit 18.3 μs 17.0 μs 1.08 ✅
track/2. L1 8sat/5K – track!-threaded 47.8 μs 47.5 μs 1.01
track/2. L1 8sat/5K – track-threaded 48.5 μs 48.4 μs 1.0
track/3. E1B 4sat/25K – track 121.0 μs 121.0 μs 1.0
track/3. E1B 4sat/25K – track! 121.0 μs 121.0 μs 0.999
track/3. E1B 4sat/25K – track! Int16 59.1 μs 59.1 μs 1.0
track/3. E1B 4sat/25K – track!-threaded 121.0 μs 121.0 μs 1.0
track/3. E1B 4sat/25K – track-threaded 121.0 μs 121.0 μs 0.999
track/4. 8L1+8E1B/25K – track 443.0 μs 443.0 μs 0.998
track/4. 8L1+8E1B/25K – track! 441.0 μs 442.0 μs 0.998
track/4. 8L1+8E1B/25K – track!-threaded 442.0 μs 442.0 μs 1.0
track/4. 8L1+8E1B/25K – track-threaded 444.0 μs 443.0 μs 1.0
track/5. multi-signal N=1/5K – track 6.19 μs 6.17 μs 1.0
track/5. multi-signal N=1/5K – track! 6.18 μs 6.24 μs 0.99
track/6. multi-signal N=2/5K – track 8.9 μs 9.02 μs 0.987
track/6. multi-signal N=2/5K – track! 9.01 μs 9.03 μs 0.998
track/7. L1 8sat/500K – track! Int16 2.28 ms 2.24 ms 1.01
track/7. L1 8sat/500K – track! OneBit 5.91 ms 1.57 ms 3.77 ✅
track/7. L1 8sat/500K – track! TwoBit 7.26 ms 2.08 ms 3.49 ✅
track/7. multi-signal N=3/5K – track 11.3 μs 11.4 μs 0.989
track/7. multi-signal N=3/5K – track! 11.9 μs 11.5 μs 1.03
track/8. L1CA presync 2 blk – track! 11.7 μs 11.6 μs 1.0
track/8. L1CA presync 20 blk – track! 113.0 μs 113.0 μs 0.994
track/8. L1CA synced 2 blk – track! 11.6 μs 11.5 μs 1.0
track/8. L1CA synced 20 blk – track! 111.0 μs 112.0 μs 0.996
time_to_load 126.0 μs 130.0 μs 0.966
Memory benchmarks (base vs PR head)

Ratio = 5f03cd2… / a0be4ef… (bytes allocated): >1 means the PR allocates less. ✅ ≥ 5 % less, ⚠️ ≥ 5 % more. /0 mark a benchmark that drops to / picks up allocations, means both revisions allocate nothing. A blank cell means the benchmark exists on only one revision (🆕 = new on the PR, 🗑 = removed).

5f03cd2 a0be4ef 5f03cd2… / a0be4ef
downconvert and correlate/CPU/Float32 2 allocs: 576 B 2 allocs: 576 B 1.0
downconvert and correlate/CPU/Float32 4ant 2 allocs: 848 B 2 allocs: 848 B 1.0
downconvert and correlate/CPU/Float64 2 allocs: 576 B 2 allocs: 576 B 1.0
downconvert and correlate/CPU/Int16 2 allocs: 576 B 2 allocs: 576 B 1.0
downconvert and correlate/CPU/Int16 4ant 2 allocs: 848 B 2 allocs: 848 B 1.0
downconvert and correlate/CPU/Int32 2 allocs: 576 B 2 allocs: 576 B 1.0
fused kernel/1-ant dynamic taps 0 allocs: 0 B 0 allocs: 0 B
fused kernel/1-ant static taps 0 allocs: 0 B 0 allocs: 0 B
fused kernel/4-ant dynamic taps 0 allocs: 0 B 0 allocs: 0 B
fused kernel/4-ant static taps 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/1-ant N=2 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/1-ant N=3 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/2-ant N=2 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/2-ant N=3 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/4-ant N=2 0 allocs: 0 B 0 allocs: 0 B
fused tuple kernel/4-ant N=3 0 allocs: 0 B 0 allocs: 0 B
track/1. Float32/2K – track 9 allocs: 944 B 9 allocs: 944 B 1.0
track/1. Float32/2K – track! 0 allocs: 0 B 0 allocs: 0 B
track/2. L1 8sat/5K – track 10 allocs: 4536 B 10 allocs: 4536 B 1.0
track/2. L1 8sat/5K – track! 0 allocs: 0 B 0 allocs: 0 B
track/2. L1 8sat/5K – track! Int16 13 allocs: 1056 B 13 allocs: 1056 B 1.0
track/2. L1 8sat/5K – track! OneBit 45 allocs: 2736 B 45 allocs: 2736 B 1.0
track/2. L1 8sat/5K – track!-threaded 0 allocs: 0 B 0 allocs: 0 B
track/2. L1 8sat/5K – track-threaded 10 allocs: 4536 B 10 allocs: 4536 B 1.0
track/3. E1B 4sat/25K – track 10 allocs: 2808 B 10 allocs: 2808 B 1.0
track/3. E1B 4sat/25K – track! 0 allocs: 0 B 0 allocs: 0 B
track/3. E1B 4sat/25K – track! Int16 5 allocs: 416 B 5 allocs: 416 B 1.0
track/3. E1B 4sat/25K – track!-threaded 0 allocs: 0 B 0 allocs: 0 B
track/3. E1B 4sat/25K – track-threaded 10 allocs: 2808 B 10 allocs: 2808 B 1.0
track/4. 8L1+8E1B/25K – track 26 allocs: 10352 B 26 allocs: 10352 B 1.0
track/4. 8L1+8E1B/25K – track! 0 allocs: 0 B 0 allocs: 0 B
track/4. 8L1+8E1B/25K – track!-threaded 0 allocs: 0 B 0 allocs: 0 B
track/4. 8L1+8E1B/25K – track-threaded 26 allocs: 10352 B 26 allocs: 10352 B 1.0
track/5. multi-signal N=1/5K – track 9 allocs: 944 B 9 allocs: 944 B 1.0
track/5. multi-signal N=1/5K – track! 0 allocs: 0 B 0 allocs: 0 B
track/6. multi-signal N=2/5K – track 9 allocs: 1408 B 9 allocs: 1408 B 1.0
track/6. multi-signal N=2/5K – track! 0 allocs: 0 B 0 allocs: 0 B
track/7. L1 8sat/500K – track! Int16 61 allocs: 25696 B 61 allocs: 25696 B 1.0
track/7. L1 8sat/500K – track! OneBit 1677 allocs: 116080 B 1677 allocs: 116080 B 1.0
track/7. L1 8sat/500K – track! TwoBit 1677 allocs: 116128 B 1677 allocs: 116128 B 1.0
track/7. multi-signal N=3/5K – track 9 allocs: 1760 B 9 allocs: 1760 B 1.0
track/7. multi-signal N=3/5K – track! 0 allocs: 0 B 0 allocs: 0 B
track/8. L1CA presync 2 blk – track! 7 allocs: 320 B 7 allocs: 320 B 1.0
track/8. L1CA presync 20 blk – track! 7 allocs: 320 B 7 allocs: 320 B 1.0
track/8. L1CA synced 2 blk – track! 7 allocs: 320 B 7 allocs: 320 B 1.0
track/8. L1CA synced 20 blk – track! 7 allocs: 320 B 7 allocs: 320 B 1.0
time_to_load 145 allocs: 11216 B 149 allocs: 11408 B 0.983

A 100 ms call spans ~100 integration steps, the post-processing-style
workload where packing the bit backends' shared band planes once per
call (instead of once per step) shows: OneBit ~2.1 -> 0.5 ms and TwoBit
~3.0 -> 0.7 ms locally, while the existing 1 ms scenarios are
unaffected by design (one step = one pack either way). Registered
twice, matching how the existing scenarios split across the two tables:

- 'track! Int16 vs Float32 / GPS L1CA, 8 sats @ 5 MHz, 100 ms buffer' —
  the PR-head backend-vs-Float32 head-to-head row;
- 'track / 7. L1 8sat/500K – track! OneBit|TwoBit' — the base-vs-head
  ratio table, so the speedup (and any future regression) shows in the
  historical comparison.

Unlike the 1 ms scenarios, these cannot use a pure-noise capture: ~100
back-to-back loop updates over noise can drift the filters to a NaN
Doppler within a single call (InexactError in CI). They track a
phase/Doppler-matched 8-PRN composite signal instead, so the loops stay
locked; verified 15-30 samples per leaf on the PR head and the pre-fix
base.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@zsoerenm
zsoerenm force-pushed the perf/pack-band-once branch from 2991178 to a0be4ef Compare July 22, 2026 12:54
@zsoerenm
zsoerenm merged commit 4fd74b9 into master Jul 22, 2026
10 checks passed
@zsoerenm
zsoerenm deleted the perf/pack-band-once branch July 22, 2026 13:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant