KV cache quantization: fp8 (KV8) + 4-bit TurboQuant (KV_TQ) on CPU, CUDA, and Metal - #399
KV cache quantization: fp8 (KV8) + 4-bit TurboQuant (KV_TQ) on CPU, CUDA, and Metal#399NeuralNotwerk wants to merge 3 commits into
Conversation
dc2bf2c to
613236f
Compare
|
Reviewed against post-v1.0.0 How the ragged branch interacts with the fp8 device shadow (does a ragged gather read the shadow rows, or does it re-upload?) is a semantic call I'm not comfortable resolving for you, and I have no CUDA hardware here to validate it. Could you rebase the series onto current |
|
Rebased onto current The ragged × fp8-shadow call, made explicit in the code (comment at the gate in Also folded in during the rebase: the batch-absorb gate keeps its |
b02bc92 to
606fe19
Compare
|
Heads-up: |
Upstream PR JustVugg#399 (KV8 fp8 / TQ4 quantized latent KV) leaves Lc/Rc NULL when a quantized tier is active — the mirror sync would deref NULL. Guard on the arrays themselves (env-independent), falling back to the CPU attention path; an fp8-aware VK mirror (upload Lc8+scale, decode e4m3 in the shader, 4x less mirror traffic) is the follow-up once JustVugg#399 merges.
CUDA validation on fleet hardware (consolidated)Ran the rebased series on our fleet box — 4× RTX 5090 (sm_120) + RTX 4090 (sm_89), EPYC 7532, 503 GB RAM, CUDA 12.8.1. (Context for the sparse CUDA numbers lately: this box pulls ~4.5 kW and North Carolina is mid-heat-wave, so it only comes up in cool windows.) Build & tests: real Correctness, live on GLM-5.2 744B ( Performance — measured at the production operating point (serve mode over HTTP, exact prod env: On the #391 heads-up ( |
|
Tested this branch on an M5 Pro, 64 GB, macOS 26.5.2, GLM-5.2 REAP-504B int4 container, expert-pruned to 168 per layer, streamed from disk. Two findings, one of which you probably want before merge.
Caveat on both: the CLI prompt is untemplated and the model degenerates into a token loop after about 30 tokens in every arm, merge-base included. The degeneration is deterministic, so the first-token divergence and the relative throughput stand, but the absolute tok/s is not a chat workload. I can rerun with different knobs or capture fuller logs if that helps narrow down the qmode question. |
Review fixes for JustVugg#399 (michael-denyer, M5 Pro): (1) f32 regression: both KV tiers OFF diverged from merge-base at the first decode token. Root cause: coli_metal_init compiled the WHOLE MTLLibrary with MTLMathModeSafe (added for the fp8 RNE encoder), silently recompiling every f32 kernel fast-math-off. A byte-identical A/B proved mm_gemv alone drifts ~4e-6, compounding to flip greedy argmax at token 1. Fix: two libraries — libFast (fast-math, all f32/MoE + quantized consumers; restores merge-base parity) and libSafe (only a_fp8enc/a_tq_enc/a_tq_enc1). fp8 stays byte-exact. (2) TQ4 ~2x slower than f32: it re-dequantized the ENTIRE cache to f32 staging every layer every token, then ran plain f32 score/clat. codec-1 (rotated int4) fix uses the Hadamard orthogonality identity q.x_hat == rotate(q).c and sum_t w_t x_hat_t == unrotate(sum_t w_t c_t): rotate the query once per head (a_tq_qrot), dot packed nibbles through the codebook (a_score_tq), accumulate (a_clat_tq), unrotate once (a_tq_unrot); parallel producer a_tq_enc1. codec-0 (PolarQuant) keeps the a_tq_deq staging path. Native vs staging: 1.4-1.9x faster, gap widens with context, equivalence ~5e-7. Tests: test_kv_tq.c gains the score/context orthogonality identities + a full MLA-consumer native-vs-dequant check; test_backend_metal.mm gains a native-vs-staging equivalence + speedup benchmark. make metal-test + test-c green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Ready for review — validated on Metal, CPU, and CUDABoth review findings are root-caused and fixed, and the KV-quant stack is now hardware-validated across all three backends (incl. an end-to-end run on GLM-5.2 744B). @michael-denyer's findings — both fixed
Validation
Also in this push
RemainingOne open item: the rebase onto the |
6c238fa to
7d62599
Compare
…— rebased onto colibri.c Rebases the JustVugg#399 series onto current dev (colibri.c + extracted headers, post-JustVugg#391), which also picks up JustVugg#445's CUDA_RELEASE_HOST pin-budget fix for free. The KV-quant work is additive — dev had no KV8/TQ code — so the split just relocated scaffolding: * colibri.c — KV8/TQ globals + KVState fp8/packed byte caches, kv_alloc, the attention consumers (CPU rotated-int4 orthogonality path; Metal/CUDA native fused kernels dispatch), env parsing, the pin_load KV8-shadow VRAM projection (merged with JustVugg#445's additive prefix budget). * kv_fp8.h / kv_tq.h — the fp8 e4m3 and rotated-int4/PolarQuant codecs (new headers). * kv_persist.h — .coli_kv disk format v2 (fp8) / v3 (TQ): magic-tagged, on-load format detection, v1->v2/v3 in-RAM upgrade + self-heal, fsync durability. * backend_metal.mm — two-library split (fixes the f32 fast-math regression: fast-math for f32/consumer kernels, safe-math only for the RNE encoders) + native codec-1 TQ4 fused attention (Hadamard-orthogonality: rotate the query once, dot packed nibbles, unrotate the context; no f32 restage). * backend_cuda.cu/.h — native codec-1 TQ4 kernel (attention_absorb_kernel_tq) + host entry. * openai_server.py — dispatcher skips unrecognized engine telemetry frames instead of tearing down (unified single/multi-slot interface, version-skew tolerant). Note: the KV-tier flag definitions (g_kv8/g_tq/g_tq_bits/g_tq_codec/g_kv_shadow) moved above the kv_persist.h include so the disk format can see them. Validated on this rebase: `make colibri` clean; `make test-c` green (incl. kv_fp8, kv_tq identities + MLA-consumer equivalence, and kv_disk v1/v2/v3 round-trip + upgrade + reject + self-heal); `make metal-test` green (f32 byte-exact, native TQ4 1.3-1.5x faster than staging, layer decode ok). CUDA compiles clean on the fleet box (sm_89+sm_120, cuda-test incl. attention_absorb_tq) and ran coherent end-to-end on GLM-5.2 744B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Rebased onto current Rebased tree is green locally: |
…— rebased onto colibri.c Rebases the JustVugg#399 series onto current dev (colibri.c + extracted headers, post-JustVugg#391), which also picks up JustVugg#445's CUDA_RELEASE_HOST pin-budget fix for free. The KV-quant work is additive — dev had no KV8/TQ code — so the split just relocated scaffolding: * colibri.c — KV8/TQ globals + KVState fp8/packed byte caches, kv_alloc, the attention consumers (CPU rotated-int4 orthogonality path; Metal/CUDA native fused kernels dispatch), env parsing, the pin_load KV8-shadow VRAM projection (merged with JustVugg#445's additive prefix budget). * kv_fp8.h / kv_tq.h — the fp8 e4m3 and rotated-int4/PolarQuant codecs (new headers). * kv_persist.h — .coli_kv disk format v2 (fp8) / v3 (TQ): magic-tagged, on-load format detection, v1->v2/v3 in-RAM upgrade + self-heal, fsync durability. * backend_metal.mm — two-library split (fixes the f32 fast-math regression: fast-math for f32/consumer kernels, safe-math only for the RNE encoders) + native codec-1 TQ4 fused attention (Hadamard-orthogonality: rotate the query once, dot packed nibbles, unrotate the context; no f32 restage). * backend_cuda.cu/.h — native codec-1 TQ4 kernel (attention_absorb_kernel_tq) + host entry. * openai_server.py — dispatcher skips unrecognized engine telemetry frames instead of tearing down (unified single/multi-slot interface, version-skew tolerant). Note: the KV-tier flag definitions (g_kv8/g_tq/g_tq_bits/g_tq_codec/g_kv_shadow) moved above the kv_persist.h include so the disk format can see them. Validated on this rebase: `make colibri` clean; `make test-c` green (incl. kv_fp8, kv_tq identities + MLA-consumer equivalence, and kv_disk v1/v2/v3 round-trip + upgrade + reject + self-heal); `make metal-test` green (f32 byte-exact, native TQ4 1.3-1.5x faster than staging, layer decode ok). CUDA compiles clean on the fleet box (sm_89+sm_120, cuda-test incl. attention_absorb_tq) and ran coherent end-to-end on GLM-5.2 744B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
7d62599 to
9b7c577
Compare
|
Rebase needed: #298 just merged into |
Upstream PR JustVugg#399 (KV8 fp8 / TQ4 quantized latent KV) leaves Lc/Rc NULL when a quantized tier is active — the mirror sync would deref NULL. Guard on the arrays themselves (env-independent), falling back to the CPU attention path; an fp8-aware VK mirror (upload Lc8+scale, decode e4m3 in the shader, 4x less mirror traffic) is the follow-up once JustVugg#399 merges.
…— rebased onto colibri.c Rebases the JustVugg#399 series onto current dev (colibri.c + extracted headers, post-JustVugg#391), which also picks up JustVugg#445's CUDA_RELEASE_HOST pin-budget fix for free. The KV-quant work is additive — dev had no KV8/TQ code — so the split just relocated scaffolding: * colibri.c — KV8/TQ globals + KVState fp8/packed byte caches, kv_alloc, the attention consumers (CPU rotated-int4 orthogonality path; Metal/CUDA native fused kernels dispatch), env parsing, the pin_load KV8-shadow VRAM projection (merged with JustVugg#445's additive prefix budget). * kv_fp8.h / kv_tq.h — the fp8 e4m3 and rotated-int4/PolarQuant codecs (new headers). * kv_persist.h — .coli_kv disk format v2 (fp8) / v3 (TQ): magic-tagged, on-load format detection, v1->v2/v3 in-RAM upgrade + self-heal, fsync durability. * backend_metal.mm — two-library split (fixes the f32 fast-math regression: fast-math for f32/consumer kernels, safe-math only for the RNE encoders) + native codec-1 TQ4 fused attention (Hadamard-orthogonality: rotate the query once, dot packed nibbles, unrotate the context; no f32 restage). * backend_cuda.cu/.h — native codec-1 TQ4 kernel (attention_absorb_kernel_tq) + host entry. * openai_server.py — dispatcher skips unrecognized engine telemetry frames instead of tearing down (unified single/multi-slot interface, version-skew tolerant). Note: the KV-tier flag definitions (g_kv8/g_tq/g_tq_bits/g_tq_codec/g_kv_shadow) moved above the kv_persist.h include so the disk format can see them. Validated on this rebase: `make colibri` clean; `make test-c` green (incl. kv_fp8, kv_tq identities + MLA-consumer equivalence, and kv_disk v1/v2/v3 round-trip + upgrade + reject + self-heal); `make metal-test` green (f32 byte-exact, native TQ4 1.3-1.5x faster than staging, layer decode ok). CUDA compiles clean on the fleet box (sm_89+sm_120, cuda-test incl. attention_absorb_tq) and ran coherent end-to-end on GLM-5.2 744B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…— rebased onto colibri.c Rebases the JustVugg#399 series onto current dev (colibri.c + extracted headers, post-JustVugg#391), which also picks up JustVugg#445's CUDA_RELEASE_HOST pin-budget fix for free. The KV-quant work is additive — dev had no KV8/TQ code — so the split just relocated scaffolding: * colibri.c — KV8/TQ globals + KVState fp8/packed byte caches, kv_alloc, the attention consumers (CPU rotated-int4 orthogonality path; Metal/CUDA native fused kernels dispatch), env parsing, the pin_load KV8-shadow VRAM projection (merged with JustVugg#445's additive prefix budget). * kv_fp8.h / kv_tq.h — the fp8 e4m3 and rotated-int4/PolarQuant codecs (new headers). * kv_persist.h — .coli_kv disk format v2 (fp8) / v3 (TQ): magic-tagged, on-load format detection, v1->v2/v3 in-RAM upgrade + self-heal, fsync durability. * backend_metal.mm — two-library split (fixes the f32 fast-math regression: fast-math for f32/consumer kernels, safe-math only for the RNE encoders) + native codec-1 TQ4 fused attention (Hadamard-orthogonality: rotate the query once, dot packed nibbles, unrotate the context; no f32 restage). * backend_cuda.cu/.h — native codec-1 TQ4 kernel (attention_absorb_kernel_tq) + host entry. * openai_server.py — dispatcher skips unrecognized engine telemetry frames instead of tearing down (unified single/multi-slot interface, version-skew tolerant). Note: the KV-tier flag definitions (g_kv8/g_tq/g_tq_bits/g_tq_codec/g_kv_shadow) moved above the kv_persist.h include so the disk format can see them. Validated on this rebase: `make colibri` clean; `make test-c` green (incl. kv_fp8, kv_tq identities + MLA-consumer equivalence, and kv_disk v1/v2/v3 round-trip + upgrade + reject + self-heal); `make metal-test` green (f32 byte-exact, native TQ4 1.3-1.5x faster than staging, layer decode ok). CUDA compiles clean on the fleet box (sm_89+sm_120, cuda-test incl. attention_absorb_tq) and ran coherent end-to-end on GLM-5.2 744B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…— rebased onto colibri.c Rebases the JustVugg#399 series onto current dev (colibri.c + extracted headers, post-JustVugg#391), which also picks up JustVugg#445's CUDA_RELEASE_HOST pin-budget fix for free. The KV-quant work is additive — dev had no KV8/TQ code — so the split just relocated scaffolding: * colibri.c — KV8/TQ globals + KVState fp8/packed byte caches, kv_alloc, the attention consumers (CPU rotated-int4 orthogonality path; Metal/CUDA native fused kernels dispatch), env parsing, the pin_load KV8-shadow VRAM projection (merged with JustVugg#445's additive prefix budget). * kv_fp8.h / kv_tq.h — the fp8 e4m3 and rotated-int4/PolarQuant codecs (new headers). * kv_persist.h — .coli_kv disk format v2 (fp8) / v3 (TQ): magic-tagged, on-load format detection, v1->v2/v3 in-RAM upgrade + self-heal, fsync durability. * backend_metal.mm — two-library split (fixes the f32 fast-math regression: fast-math for f32/consumer kernels, safe-math only for the RNE encoders) + native codec-1 TQ4 fused attention (Hadamard-orthogonality: rotate the query once, dot packed nibbles, unrotate the context; no f32 restage). * backend_cuda.cu/.h — native codec-1 TQ4 kernel (attention_absorb_kernel_tq) + host entry. * openai_server.py — dispatcher skips unrecognized engine telemetry frames instead of tearing down (unified single/multi-slot interface, version-skew tolerant). Note: the KV-tier flag definitions (g_kv8/g_tq/g_tq_bits/g_tq_codec/g_kv_shadow) moved above the kv_persist.h include so the disk format can see them. Validated on this rebase: `make colibri` clean; `make test-c` green (incl. kv_fp8, kv_tq identities + MLA-consumer equivalence, and kv_disk v1/v2/v3 round-trip + upgrade + reject + self-heal); `make metal-test` green (f32 byte-exact, native TQ4 1.3-1.5x faster than staging, layer decode ok). CUDA compiles clean on the fleet box (sm_89+sm_120, cuda-test incl. attention_absorb_tq) and ran coherent end-to-end on GLM-5.2 744B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
9b7c577 to
c32e38f
Compare
|
Status: this conflicts with On the substance: KV quantization is a real lever here and I'd like it, but the shape of the PR is what's slowing it down. It's ~3,400 lines across 20 files covering two schemes (fp8 KV8 and 4-bit TurboQuant) on three backends (CPU, CUDA, Metal). I can validate the CPU path; I have neither NVIDIA nor Apple hardware, so the other two thirds rest entirely on your measurements. That's a lot to accept in one merge. What would unblock it: land the CPU path first, on its own. It's the one I can verify end to end against the oracle, and once the format and the KV plumbing are settled and merged, the CUDA and Metal kernels become much smaller, more focused follow-ups — each with a hardware owner who can gate it (this project's rule for GPU changes is a token-identity check from someone with the silicon). Also worth stating explicitly in the PR: what happens to KV persistence ( Tell me if splitting is workable and I'll prioritise the CPU half. |
Upstream PR JustVugg#399 (KV8 fp8 / TQ4 quantized latent KV) leaves Lc/Rc NULL when a quantized tier is active — the mirror sync would deref NULL. Guard on the arrays themselves (env-independent), falling back to the CPU attention path; an fp8-aware VK mirror (upload Lc8+scale, decode e4m3 in the shader, 4x less mirror traffic) is the follow-up once JustVugg#399 merges.
|
Splitting is workable — I'll carve the CPU path out as its own PR on current Teacher-forcing token-exact vs the transformers oracle
Greedy generation (decode path — this is where the native CUDA/Metal quantized-KV kernels engage):
The TQ4 generated sequence is byte-identical CUDA vs Metal vs CPU. Reading this honestly:
KV persistence across mode changesAlready handled the way you describe, and covered by The test covers v1→v2/v3 upgrade (old file untouched until the first save rewrites it), v2/v3-under-f32 reject, TQ codec/bits mismatch reject, and the append-side self-heal. Perf reference (already posted above, environment for the record)Resident 744B, driver 610.43.02, CUDA 12.8.1 runtime, 4×RTX 5090 + RTX 4090 ( Next from me: the CPU-only PR (KV8 + KV_TQ consumers, |
|
Split done — the CPU path is now its own PR: #553 (1,282 lines / 12 files, no backend code), rebased on today's Converting this PR to draft: it stays as the reference for the full three-backend implementation, and once #553 lands I'll rebase the CUDA and Metal kernels out of it as the two focused follow-ups (each with its hardware validation + token-identity data — the CUDA half is already validated on 4×5090+4090 at f32-parity decode on the 744B). |
…— rebased onto colibri.c Rebases the JustVugg#399 series onto current dev (colibri.c + extracted headers, post-JustVugg#391), which also picks up JustVugg#445's CUDA_RELEASE_HOST pin-budget fix for free. The KV-quant work is additive — dev had no KV8/TQ code — so the split just relocated scaffolding: * colibri.c — KV8/TQ globals + KVState fp8/packed byte caches, kv_alloc, the attention consumers (CPU rotated-int4 orthogonality path; Metal/CUDA native fused kernels dispatch), env parsing, the pin_load KV8-shadow VRAM projection (merged with JustVugg#445's additive prefix budget). * kv_fp8.h / kv_tq.h — the fp8 e4m3 and rotated-int4/PolarQuant codecs (new headers). * kv_persist.h — .coli_kv disk format v2 (fp8) / v3 (TQ): magic-tagged, on-load format detection, v1->v2/v3 in-RAM upgrade + self-heal, fsync durability. * backend_metal.mm — two-library split (fixes the f32 fast-math regression: fast-math for f32/consumer kernels, safe-math only for the RNE encoders) + native codec-1 TQ4 fused attention (Hadamard-orthogonality: rotate the query once, dot packed nibbles, unrotate the context; no f32 restage). * backend_cuda.cu/.h — native codec-1 TQ4 kernel (attention_absorb_kernel_tq) + host entry. * openai_server.py — dispatcher skips unrecognized engine telemetry frames instead of tearing down (unified single/multi-slot interface, version-skew tolerant). Note: the KV-tier flag definitions (g_kv8/g_tq/g_tq_bits/g_tq_codec/g_kv_shadow) moved above the kv_persist.h include so the disk format can see them. Validated on this rebase: `make colibri` clean; `make test-c` green (incl. kv_fp8, kv_tq identities + MLA-consumer equivalence, and kv_disk v1/v2/v3 round-trip + upgrade + reject + self-heal); `make metal-test` green (f32 byte-exact, native TQ4 1.3-1.5x faster than staging, layer decode ok). CUDA compiles clean on the fleet box (sm_89+sm_120, cuda-test incl. attention_absorb_tq) and ran coherent end-to-end on GLM-5.2 744B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
c32e38f to
23b7628
Compare
|
Rebased onto current Four files conflicted. Two were mechanical, three were semantic and worth review: 1. DSA gather — dev's 2. 3. HIP portability — this one is a real behavior change. 4. Validation on the rebased treeRun here on macOS 26.5.1 / Apple M4 / 32 GB:
What this does not re-validateThe fleet-box CUDA numbers I posted on the 20th were measured on the pre-split tree, and the DSA-gather routing and the reserve accounting both changed in this rebase. Those need a re-run on |
…— rebased onto colibri.c Rebases the JustVugg#399 series onto current dev (colibri.c + extracted headers, post-JustVugg#391), which also picks up JustVugg#445's CUDA_RELEASE_HOST pin-budget fix for free. The KV-quant work is additive — dev had no KV8/TQ code — so the split just relocated scaffolding: * colibri.c — KV8/TQ globals + KVState fp8/packed byte caches, kv_alloc, the attention consumers (CPU rotated-int4 orthogonality path; Metal/CUDA native fused kernels dispatch), env parsing, the pin_load KV8-shadow VRAM projection (merged with JustVugg#445's additive prefix budget). * kv_fp8.h / kv_tq.h — the fp8 e4m3 and rotated-int4/PolarQuant codecs (new headers). * kv_persist.h — .coli_kv disk format v2 (fp8) / v3 (TQ): magic-tagged, on-load format detection, v1->v2/v3 in-RAM upgrade + self-heal, fsync durability. * backend_metal.mm — two-library split (fixes the f32 fast-math regression: fast-math for f32/consumer kernels, safe-math only for the RNE encoders) + native codec-1 TQ4 fused attention (Hadamard-orthogonality: rotate the query once, dot packed nibbles, unrotate the context; no f32 restage). * backend_cuda.cu/.h — native codec-1 TQ4 kernel (attention_absorb_kernel_tq) + host entry. * openai_server.py — dispatcher skips unrecognized engine telemetry frames instead of tearing down (unified single/multi-slot interface, version-skew tolerant). Note: the KV-tier flag definitions (g_kv8/g_tq/g_tq_bits/g_tq_codec/g_kv_shadow) moved above the kv_persist.h include so the disk format can see them. Validated on this rebase: `make colibri` clean; `make test-c` green (incl. kv_fp8, kv_tq identities + MLA-consumer equivalence, and kv_disk v1/v2/v3 round-trip + upgrade + reject + self-heal); `make metal-test` green (f32 byte-exact, native TQ4 1.3-1.5x faster than staging, layer decode ok). CUDA compiles clean on the fleet box (sm_89+sm_120, cuda-test incl. attention_absorb_tq) and ran coherent end-to-end on GLM-5.2 744B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
23b7628 to
85852e2
Compare
|
Rebased onto current The rebaseOne conflict, in Also folded in from the #553 review pass (cherry-picked, so it collapses cleanly when #553 lands):
Token identity across four backends, on this treeThe claim this PR has been making is that the quantized flips are deterministic quantization loss rather than backend divergence. That now holds across two ISAs and two GPU vendors, all re-measured on
Identical positions, identical tokens, four backends. The AMD column is new — these kernels had never been run on ROCm. Other backends on the rebased tree:
What this still does not claimNo perf numbers are carried forward. The fleet-box CUDA throughput figures I posted on the 20th were measured pre-split, and this rebase changed both the DSA-gather routing and the One observation for whoever owns the AMD path: on this box the tiny-model teacher-forcing rate went from ~3230 pos/s to ~159 pos/s after #599 enabled rocWMMA. It's a 5-layer random-weight model where launch overhead dominates, so it may well be meaningless — but it's a 20× delta on |
Upstream PR JustVugg#399 (KV8 fp8 / TQ4 quantized latent KV) leaves Lc/Rc NULL when a quantized tier is active — the mirror sync would deref NULL. Guard on the arrays themselves (env-independent), falling back to the CPU attention path; an fp8-aware VK mirror (upload Lc8+scale, decode e4m3 in the shader, 4x less mirror traffic) is the follow-up once JustVugg#399 merges.
…— rebased onto colibri.c Rebases the JustVugg#399 series onto current dev (colibri.c + extracted headers, post-JustVugg#391), which also picks up JustVugg#445's CUDA_RELEASE_HOST pin-budget fix for free. The KV-quant work is additive — dev had no KV8/TQ code — so the split just relocated scaffolding: * colibri.c — KV8/TQ globals + KVState fp8/packed byte caches, kv_alloc, the attention consumers (CPU rotated-int4 orthogonality path; Metal/CUDA native fused kernels dispatch), env parsing, the pin_load KV8-shadow VRAM projection (merged with JustVugg#445's additive prefix budget). * kv_fp8.h / kv_tq.h — the fp8 e4m3 and rotated-int4/PolarQuant codecs (new headers). * kv_persist.h — .coli_kv disk format v2 (fp8) / v3 (TQ): magic-tagged, on-load format detection, v1->v2/v3 in-RAM upgrade + self-heal, fsync durability. * backend_metal.mm — two-library split (fixes the f32 fast-math regression: fast-math for f32/consumer kernels, safe-math only for the RNE encoders) + native codec-1 TQ4 fused attention (Hadamard-orthogonality: rotate the query once, dot packed nibbles, unrotate the context; no f32 restage). * backend_cuda.cu/.h — native codec-1 TQ4 kernel (attention_absorb_kernel_tq) + host entry. * openai_server.py — dispatcher skips unrecognized engine telemetry frames instead of tearing down (unified single/multi-slot interface, version-skew tolerant). Note: the KV-tier flag definitions (g_kv8/g_tq/g_tq_bits/g_tq_codec/g_kv_shadow) moved above the kv_persist.h include so the disk format can see them. Validated on this rebase: `make colibri` clean; `make test-c` green (incl. kv_fp8, kv_tq identities + MLA-consumer equivalence, and kv_disk v1/v2/v3 round-trip + upgrade + reject + self-heal); `make metal-test` green (f32 byte-exact, native TQ4 1.3-1.5x faster than staging, layer decode ok). CUDA compiles clean on the fleet box (sm_89+sm_120, cuda-test incl. attention_absorb_tq) and ran coherent end-to-end on GLM-5.2 744B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
85852e2 to
d800ace
Compare
Upstream PR JustVugg#399 (KV8 fp8 / TQ4 quantized latent KV) leaves Lc/Rc NULL when a quantized tier is active — the mirror sync would deref NULL. Guard on the arrays themselves (env-independent), falling back to the CPU attention path; an fp8-aware VK mirror (upload Lc8+scale, decode e4m3 in the shader, 4x less mirror traffic) is the follow-up once JustVugg#399 merges.
…— rebased onto colibri.c Rebases the JustVugg#399 series onto current dev (colibri.c + extracted headers, post-JustVugg#391), which also picks up JustVugg#445's CUDA_RELEASE_HOST pin-budget fix for free. The KV-quant work is additive — dev had no KV8/TQ code — so the split just relocated scaffolding: * colibri.c — KV8/TQ globals + KVState fp8/packed byte caches, kv_alloc, the attention consumers (CPU rotated-int4 orthogonality path; Metal/CUDA native fused kernels dispatch), env parsing, the pin_load KV8-shadow VRAM projection (merged with JustVugg#445's additive prefix budget). * kv_fp8.h / kv_tq.h — the fp8 e4m3 and rotated-int4/PolarQuant codecs (new headers). * kv_persist.h — .coli_kv disk format v2 (fp8) / v3 (TQ): magic-tagged, on-load format detection, v1->v2/v3 in-RAM upgrade + self-heal, fsync durability. * backend_metal.mm — two-library split (fixes the f32 fast-math regression: fast-math for f32/consumer kernels, safe-math only for the RNE encoders) + native codec-1 TQ4 fused attention (Hadamard-orthogonality: rotate the query once, dot packed nibbles, unrotate the context; no f32 restage). * backend_cuda.cu/.h — native codec-1 TQ4 kernel (attention_absorb_kernel_tq) + host entry. * openai_server.py — dispatcher skips unrecognized engine telemetry frames instead of tearing down (unified single/multi-slot interface, version-skew tolerant). Note: the KV-tier flag definitions (g_kv8/g_tq/g_tq_bits/g_tq_codec/g_kv_shadow) moved above the kv_persist.h include so the disk format can see them. Validated on this rebase: `make colibri` clean; `make test-c` green (incl. kv_fp8, kv_tq identities + MLA-consumer equivalence, and kv_disk v1/v2/v3 round-trip + upgrade + reject + self-heal); `make metal-test` green (f32 byte-exact, native TQ4 1.3-1.5x faster than staging, layer decode ok). CUDA compiles clean on the fleet box (sm_89+sm_120, cuda-test incl. attention_absorb_tq) and ran coherent end-to-end on GLM-5.2 744B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add opt-in (PROF=1) instrumentation for expert-weight disk I/O rates, which were previously reported only as seconds, never as throughput: - pin_load: '[PROF] pin load: X GB read in Ys = Z GB/s aggregate | N experts @ M experts/s' — the initial hot-expert load rate off disk at startup. Gated on getenv(PROF) directly since g_prof is parsed after pin_load runs. - prof_report: '[PROF] disk stream: E experts/s | A GB/s aggregate over the phase (Cx avg read concurrency, P GB/s per loader thread)' — the live streaming rate during prefill/decode. Aggregate = bytes/wall; read-service seconds sum across parallel loaders, so io_svc/wall is the concurrency and bytes/io_svc the per-thread rate (aggregate = per-thread x concurrency). Additive only; with PROF unset every mode's output stays byte-identical. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two no-silent-error gaps in the KV_TQ env handling, both found in review. 1. Power-of-two guard. Both codecs rotate through a radix-2 FWHT, and coli_kvq_quant_row returns an inert radius 0 for any other width. On a model whose kv_lora/qk_rope are not powers of two, that meant EVERY latent row quantized to zero and the engine generated confident garbage with no diagnostic -- the same silent-misread class the .coli_kv tier magic exists to prevent, just reached through model shape instead of file format. Now checked once after model_init and refused with the shapes named and KV8 (no width constraint) suggested. GLM-5.2 is 512/64, so nothing that works today changes; this only fires where the codec cannot represent the model. test_kv_tq pins the underlying behavior for all four entry points (both codecs, both dispatch paths) so the guard cannot become quietly harmless. 2. KV_TQ=1 clamped UP to 2, handing "just turn it on" the most aggressive, lowest-quality tier. It now lands on the recommended 4-bit tier and says so. >6 still clamps down to the grid. Unchanged: default (no KV env) is 32/32 token-exact, KV8 30/32, KV_TQ=4 23/32 on the tiny oracle -- identical to the pre-change numbers.
d800ace to
0927c45
Compare
Upstream PR JustVugg#399 (KV8 fp8 / TQ4 quantized latent KV) leaves Lc/Rc NULL when a quantized tier is active — the mirror sync would deref NULL. Guard on the arrays themselves (env-independent), falling back to the CPU attention path; an fp8-aware VK mirror (upload Lc8+scale, decode e4m3 in the shader, 4x less mirror traffic) is the follow-up once JustVugg#399 merges.
|
Status check, and a question rather than a rebase request.
As I read it:
If that is right, rebasing both is wasted effort for you: they will keep conflicting with each other as well as with If I have the relationship backwards, tell me and I will treat #399 as the live one instead. Either way, one thing changed on our side that makes the next rebase cheaper: |
|
Ampere data point: this branch builds and runs on 4× RTX A6000 (sm_86) with CUDA 13, on GLM-5.2 744B. Posting it because @JustVugg's sequencing note says the CUDA follow-up would be gated on a hardware owner's check, and this is the hardware. It is not the token-identity check — I explain below exactly why I cannot claim that one, so nobody treats this as more than it is. Host
Buildmake colibri CUDA=1 CUDA_ARCH=sm_86 CUDA_HOME=~/cuda13Exit 0. Two warnings, one of them yours-adjacent and harmless: Binary links RunsThree conditions,
No crash, no Generated text, prompt "Explique en une phrase ce qu'est un mixture of experts": All three fluent, correct, and semantically equivalent. What I did not establish, and why it matters hereThe three outputs are not token-identical despite
So the honest statement is: on this host I cannot currently perform a token-identity check, because I have not shown the baseline is token-identical to itself. That seems worth knowing before the CUDA follow-up is gated on exactly that test, since if multi-GPU is non-deterministic the gate needs to be a distributional check rather than an equality one. The host is powered down for now. When it is back the first thing I will run is baseline ×3 with the same seed and config, and report whether identity holds at all. If it does, the KV8 comparison becomes meaningful and I will run it properly. Unrelated context that argues for landing thisPer-phase profile from the baseline run above, same host, decode: Attention dominates decode by a wide margin on this class of host, so KV cache width is not a memory footnote here — it is on the critical path. And on the capacity side, Happy to run whatever specific check you want when the machine is back up — including on #553's CPU path, where determinism should be much easier to establish. |
|
Follow-up to my previous comment: the prerequisite is answered, and it came out the way that helps you. I said I could not perform the token-identity check because I had never shown the baseline was token-identical to itself on 4 GPUs — and that if multi-GPU reduction order drifted, the gate you proposed would need to be distributional rather than an equality. Ran it. Six identical requests, GLM-5.2 744B, 4× RTX A6000, Deterministic to the token on 4 GPUs. My worry about cross-device FP reduction order at So the check you wanted to gate the CUDA follow-up on is achievable on this hardware, and it can stay an equality test rather than becoming a distributional one. That is the simpler gate and the stronger one. What this means for my earlier numbersIt also removes the ambiguity I flagged. In my first comment the three conditions produced different text: I refused to read that as "KV8 changes the output" because I could not rule out run-to-run noise. I now can: there is no run-to-run noise. Those three outputs differ because something in the configuration differed — not because the engine wanders. I am still not claiming it is KV8, for one concrete reason: That is a real question and it now has a clean way to be answered, which it did not have before. The proper run is: baseline ×2 (confirm identity within the PR build, not just the production build), then The host is running other measurements right now. I will post that diff. One correction to my earlier comment, since it affects how you read the profile I postedI attached a per-phase profile claiming attention is 64 % of decode. Fresh runs at the same The number in that profile that survives and is worth your eye is different, and it is from 19.67 GB/s for the CPU-side routed expert path, on 8-channel DDR4-2400 that sustains ~85. That is 23 % of the host's memory bandwidth, and it caps this machine at ~3 tok/s on its own regardless of anything else. I am sweeping |
…a refuted hypothesis
The number that matters most in this whole report turns out to be one line
of PROF=1 output:
P0-EXEC: routed CPU 10.967s / 19.67 GB/s (10158 row)
19.67 GB/s on a host whose 8-channel DDR4-2400 sustains ~85. The CPU-side
routed expert path uses 23% of available memory bandwidth - the quantified
form of "4-5 cores of 24 busy" already in §5. It caps the machine on its
own:
12.1 GB of expert weights touched per token, 6.6 GB non-resident
at 19.67 GB/s -> 333 ms/token -> 3.0 tok/s (today)
at 85 GB/s -> 77 ms/token -> 13.0 tok/s (hard ceiling)
Everything else in this report is bounded above by those two lines.
§5 also gains a caveat it needed. The "attention is 64% of decode" profile
was a ~750-token generation; at 64 tokens the ranking inverts to
expert-matmul 46% / attention 26%. The share is generation-length
dependent. Both are kept - the dependence is the finding, not either point.
§10 (new): COLI_CUDA_ATTN_SHARD / COLI_CUDA_PIPE_SHARD measured separately
for the first time. S0 1.97, ATTN 2.05, PIPE 2.06, both 2.15 - +9.1%, not
the +49% I attributed to them. That gap was an artefact of comparing a
~750-token run against a 64-token one. Stated as the refutation it is.
§11 (new): six identical requests at temperature 0 produce one distinct
output on 4 GPUs. Token-identical, so byte-for-byte diffing is a usable
acceptance test for CUDA-path changes (JustVugg#399).
§12 (new): against llama.cpp on the same model, same SSD model on both
sides. llama.cpp -cmoe: prefill 239 tok/s, decode 1.58, VRAM 26 GB.
colibri: 198, 1.56, VRAM 176 GB. llama.cpp matches decode and beats
prefill by 20% while leaving the GPUs almost unused - and that is its
least favourable configuration. Not a like-for-like architectural
comparison (GGUF fuses a layer's experts, so llama.cpp cannot place them),
but on this host the per-expert placement machinery is not currently
buying anything over keeping experts in RAM. Consistent with the 19.67
GB/s ceiling, which binds both engines.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Summary
Quantized KV cache for the MLA latent across all three backends, as opt-in tiers:
KV8=1— fp8 e4m3 + per-row scale (~3.9× less KV RAM than f32), CPU + CUDA + Metal.KV_TQ=4— 4-bit KV (~7.8× less): rotated-int4 + Lloyd-Max codebook by default (full reconstruction, so it corrects the MLA key and value), or PolarQuant (arXiv 2502.02617,KV_TQ_POLAR=1, variable 2-6 bit). CPU + Metal; under CUDA the quantized paths self-disable (guarded, falls back to CPU attention).COLI_METAL=0to force CPU) — Apple unified memory has no separate VRAM pool to overflow, so the CUDA-style opt-in gate doesn't apply.Supersedes #299/#300, which I withdrew for fleet soak-testing: the VRAM-budget interaction I flagged there (fp8 shadow vs
CUDA_EXPERT_GB=auto) is fixed in this series ("KV8 shadow-aware auto VRAM budget" + "KV_SHADOW auto keyed on ragged decode"), and the stack has since gathered the review fixes, vectorized fp8 kernels (15.8× prefill, 4-6× decode vs scalar), the 4-bit tier, and a second architecture (Metal) worth of validation.Commits (bisectable, each builds)
.coli_kvv2 — encoder bit-exact with the engine's consumers (RNE, per-row amax/448 scale); disk format v2 with quantize-on-upgrade from v1, reject-on-mismatch, self-heal.kv_tq.hcodecs (randomized-Hadamard rotation + recursive polar / rotated-int4 + Lloyd codebook), MSL kernelsa_fp8enc/a_score8/a_clat8/a_tq_enc/a_tq_deq, 3-wayqmodein the fused layer decode,.coli_kvv3 (COLIKV3,h[7]=codec<<8|bits), Metal opt-out default, tests.Why rotated-int4 for the 4-bit tier
MLA's latent is both key and value, so score-only fixes (QJL-style) leave the value error uncorrected. Full reconstruction through a randomized-Hadamard rotation + fixed 16-level Lloyd-Max codebook (N(0,1)-optimal, scaled by radius/√n) measures ~0.096 rel-L2 at 4 bits vs ~0.147 for recursive-polar — and the rotation makes the radius invariant, so one per-row scale survives. 4-bit KV is a quality trade (tiny-oracle: f32 32/32, KV8 30/32, TQ4-int4 23/32) — that's the physics of 4 bits against MLA, not codec slack; KV8 is the lossless-feeling tier, TQ4 the memory-constrained one.
Validation
make test-cgreen (new suites:test_kv_fp8exhaustive e4m3 roundtrip + RNE;test_kv_tqrotation self-inverse, roundtrip, distortion, inert rows;test_kv_diskv1/v2/v3 round-trip + upgrade/reject/self-heal;test_kv_alloc).make metal-testgreen — fp8 GPU encoder byte-exact vscoli_fp8_enc; TQ GPU↔CPU roundtrip ~3e-8; fused attention parity vs CPU ~5e-6 for fp8/int4/polar; full-layer residual parity for f32/KV8/TQ (~4-5e-6).DRAFT=2decodes at 0.30 tok/s with 100% draft acceptance on a counting prompt (MTP head repaired per tools: repair int4-converted MTP heads in place; warn on --mtp --ebits <8 #397 — the two features compose).!g_tqguards on the absorb gates — no TQ CUDA kernels exist yet, phase 2). Not re-compiled on this macOS machine; last CUDA build/test was the fleet run.Notes for review
KV8/KV_TQunset → byte-identical behavior; the quantized byte-caches are NULL and every consumer branches ong_kv8||g_tq).COLI_METAL=0) and non-Metal builds behave exactly as the CPU tier.🤖 Generated with Claude Code