Skip to content

Commit ed88f21

Browse files
committed
Add fail-closed GDN BA trace gate
FOLLOWING_AGENTS_PROTOCOL Assisted-by: Codex:GPT-5 [Codex]
1 parent 5e1b1cc commit ed88f21

18 files changed

Lines changed: 971 additions & 99 deletions

.agents/coordination.md

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -113,7 +113,7 @@ time owns the GB10. Results without the lock for their entire run are discarded.
113113
| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
114114
|---|---|---|---|---|---|---|---|
115115
| `CLAIM-PR3` | `KERNEL-GDN-AOT-BF16`, `KERNEL-GDN-SCRATCH` | root takeover of stopped `validate_pr3` / `complete_pr3` stream | primary recovery tree `/home/mudler/_git/vllm.cpp-pr3-validate`; 27B default/component integration in `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; DGX evidence `~/work/vllm.cpp-noPy` plus `~/work/vllm.cpp-nvfp4-small-m/debug/gdn-out-bf16-c16-ab-20260712` | recovery branch integrated into `main` by `a767188`; current checkpoint on `codex/nvfp4-small-m` | PR #3 files and rows `KERNEL-GDN-AOT-BF16`+`KERNEL-GDN-SCRATCH`; 27B-only `GdnOutDType` default/f32 override in `qwen3_5.cpp`; ledger/inventory/roadmap evidence. No 35B default change. GPU lock: residual trace/pool classification remains separate from FP4 W3 | `ACTIVE` | 2026-07-13 (vendored BF16 H32/H48 AOT/safety evidence and native 16/16 correctness remain green. Binding `3f256ab` has c16 total at 1.027889× but mean TPOT/ITL at 0.987450×. The dynamic scan ranks packed pure-decode GDN fusion after W3-C removes tactic-selection confounding. All 35B paths keep f32) |
116-
| `CLAIM-SERVE-GATE-1` | `SERVE-GATE-ONLINE`, `KERNEL-GEMM-NVFP4-W4A4`, `KERNEL-ATTN-FA2` | root takeover | binding grid `~/work/vllm.cpp-online-gate/evidence/3f256abdbb558e162bf8a2196284deb119648560`; active worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; finalized c2 root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; W1 immutable root `~/work/vllm.cpp-gdn-ba/immutable-581d335fec2e5a96d9ccbb38c1ec001c39ac1789` | current checkpoint `codex/nvfp4-small-m` | Online-gate execution and harness work, including immutable evidence for the now-unclaimed `GATING` W1 BA implementation. qkvz code, loader-memory repair, exact grid, and 35B performance are excluded. Any A/B or trace series uses one `flock`; further `KERNEL-GEMM-BF16` production code requires a fresh `ACTIVE` claim | `ACTIVE` | 2026-07-14 (W1 implementation claim released at `GATING`. Clean pushed `581d335` closes F32-output core correctness/safety: merged/split 27B **235/235 + 16/16**, strict memcheck **590/590**, native 35B **315/315**. BF16 output fails **233/235**; exact-c2 merged/split harness, trace/component and rounding repair remain pending; binding 55/124, no speed credit) |
116+
| `CLAIM-SERVE-GATE-1` | `SERVE-GATE-ONLINE`, `KERNEL-GEMM-NVFP4-W4A4`, `KERNEL-ATTN-FA2` | root takeover | binding grid `~/work/vllm.cpp-online-gate/evidence/3f256abdbb558e162bf8a2196284deb119648560`; active worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; finalized c2 root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; W1 safety root `~/work/vllm.cpp-gdn-ba/immutable-581d335fec2e5a96d9ccbb38c1ec001c39ac1789`; next trace root `~/work/vllm.cpp-gdn-ba-trace/<pushed-sha>` | current checkpoint `codex/nvfp4-small-m` | Online-gate execution and harness work, including immutable evidence for the now-unclaimed `GATING` W1 BA implementation. qkvz code, loader-memory repair, exact grid, and 35B performance are excluded. Any A/B or trace series uses one `flock`; further `KERNEL-GEMM-BF16` production code requires a fresh `ACTIVE` claim | `ACTIVE` | 2026-07-14 (W1 core correctness/safety remains closed at `581d335`; BF16 output fails **233/235**. Explicit merged/split c2 contracts, one-lock paired driver, provenance validation and completion-marker-last finalizer are locally green at **68/68** tool tests. Run/finalize the pushed SHA to prove 963/145 versus 1,011/193 before component work; binding 55/124, no speed credit) |
117117
| `CLAIM-NVFP4-SMALL-M-2` | `KERNEL-GEMM-NVFP4-W4A4` (`W2`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; immutable source/evidence `~/work/vllm.cpp-nvfp4-small-m/b5c6e4fd65cdacea8f378e18ae101ebf521e8f01/w2` plus exact online gate | `codex/nvfp4-small-m` | W2 only: 32 tactics, merged CT gate/up and one-input `SiluAndMulFp4Quant`, with independent fallbacks and full correctness/safety/component/oracle gates | `RELEASED` | 2026-07-12 (implementation/correctness complete; strict old-oracle acceptance failed. The vLLM 0.24.0 ratios are historical diagnostics; W3 remains the trace-driven repair track under the new v0.25.0 denominator) |
118118
| `CLAIM-NVFP4-SMALL-M-3` | `KERNEL-GEMM-NVFP4-W4A4` (`W3-C`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; oracle cache fixture under `tests/fixtures/nvfp4_flashinfer_v025_gb10/`; immutable C3/C3R/corrected component `~/work/vllm.cpp-nvfp4-persistent/d211b8f80fff831a712f0bfafa4f65f1abe1892d/evidence` | `codex/nvfp4-small-m` | W3-C document/import/atomicity, ready-map/lifecycle/5,000-us wiring and corrected same-plan gates only; other levers excluded | `RELEASED` | 2026-07-13 (W3-C reproduction control complete: six-process and corrected 12-leg components use identical 64/64 maps with zero tuning/misses. W3-E strict-fails 39/40 timing + 1/8 memory; no exact grid/35B performance) |
119119
| `CLAIM-NVFP4-SMALL-M-4` | `KERNEL-GEMM-NVFP4-W4A4` (`W3-F`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; immutable source/evidence `~/work/vllm.cpp-nvfp4-device-alpha/7517af4f983fe322ac88ce2d9869e1441b7be3fd/` | `codex/nvfp4-small-m` | W3-F only: model-owned device F32 alpha, tensor op ABI/validation, `VT_FP4_DEVICE_ALPHA=0` fallback, ported tests, safety/model/node-trace and frozen-plan c2/c16 gates per [spike](specs/nvfp4-device-alpha.md). Quant/GDN/attention/host-weight changes excluded | `RELEASED` | 2026-07-13 (local/CUDA/operator/memcheck/model/trace gates pass. Completed 12-leg/612-request component is c2/c16 1.001967×/1.000144× but strict-fails 27/40 timing + 3/8 memory. No speed credit/exact grid/35B performance; the completed scan moves W3-G FA2 decode under CLAIM-SERVE-GATE-1) |
@@ -132,7 +132,7 @@ open. qkvz, exact-grid and 35B work stay excluded until those gates close.
132132

133133
| Priority | Row/block | Dependency | Next handoff | State |
134134
|---|---|---|---|---|
135-
| 1 | `KERNEL-GEMM-BF16` merged GDN projections for `SERVE-GATE-ONLINE` | `3f256ab` remains 55/124; pushed `581d335` closes F32-output core correctness/safety, BF16 output fails, and no trace/component speed credit exists | extend exact-c2 validation/driver for default and split; prove 145-vs-193 graph structure, resolve BF16 rounding, then complete c2/c16 before qkvz | `GATING` |
135+
| 1 | `KERNEL-GEMM-BF16` merged GDN projections for `SERVE-GATE-ONLINE` | `3f256ab` remains 55/124; pushed `581d335` closes F32-output core correctness/safety, BF16 output fails, and the mode-aware trace harness is locally green with no GPU/speed credit | run/finalize its pushed-SHA merged/split series to prove 145-vs-193 graph structure, resolve BF16 rounding, then complete c2/c16 before qkvz | `GATING` |
136136
| 2 | `SERVE-ASYNC-LLM` block (`ENG-CORE-BUSY-LOOP`, `SERVE-ASYNC-LLM`, `ENG-ASYNC-SCHED`, `ENG-PRIORITY-SCHED`) | joint spike accepted ([async-serving.md](specs/async-serving.md)); W1/W2/W4 implemented/`GATING`, W3 `READY`; complete control shows W3 is neutral for order-0 speed | retain W3 as unclaimed parity work until order 0 closes; do not fold it into the active kernel claim | `GATING` |
137137
| 3 | `SERVE-E2E-NIGHTLY` | `SERVE-GATE-ONLINE` evidence where benchmarks overlap | write spike and CI/nightly split | `INVENTORIED` |
138138
| 4 | C1 kernel drop-in alignment | accepted kernel-family inventory + [drop-in ABI spike](specs/dropin-kernel-abi.md) | `BACKEND-ABI-VT` W0 is implemented and CPU-green; run cross-build + flock-held GB10 runtime/capture/model/trace handoff, close scalar-forwarder/backend-shim debts, then migrate one family per independently gated checkpoint | `GATING` |

.agents/engine-matrix.md

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -20,8 +20,10 @@ vLLM qkvz+ba across 48 GDN layers. The accepted
2020
[merged-projection spike](specs/gdn-merged-input-projections.md) makes
2121
`KERNEL-GEMM-BF16` W1 `GATING`. Clean pushed `581d335` closes the one-owner
2222
F32-output BA core correctness/safety gate; BF16 output fails the token near-tie,
23-
while exact trace/component evidence remains pending. qkvz stays excluded.
24-
No trace duration earns speed credit. Host PSS/RSS separately retains a
23+
while exact trace/component evidence remains pending. The explicit mode-aware
24+
c2 driver/contracts/finalizer pass 68/68 local tool tests; immutable GPU
25+
execution is next. qkvz stays excluded. No trace duration earns speed credit.
26+
Host PSS/RSS separately retains a
2527
**22.920 GiB** CPU weight mirror plus source mmap residency.
2628

2729
| Area | Rows | `ANCHOR-BACKFILL` | `PARTIAL` | `SPIKE` | `READY` | `ACTIVE` | `GATING` | `INVENTORIED` |
@@ -157,7 +159,7 @@ claims it.
157159
| `SERVE-C-ABI` | Stable LocalAI-style C FFI (17 exported symbols; blocking and nonblocking request handles) | T0 | Original project ABI; pinned vLLM has no C ABI | `include/vllm.h:143,181,207`; `src/capi/vllm_c.cpp:229,264,327,391` | `tests/capi/test_capi.cpp:320,428,505,574,606,640`; `tests/capi/test_dlopen.cpp:77,86`; `tests/capi/c_header_compile.c:1` | `planned: specs/c-api-library.md` | `ANCHOR-BACKFILL` | - |
158160
| `SERVE-CPP-API` | Rich `LLM` and `AsyncLLM` C++ API | T1 | `vllm/entrypoints/llm.py:66,422`; `vllm/v1/engine/async_llm.py:70` | - | - | `planned: specs/cpp-api.md` | `INVENTORIED` | - |
159161
| `SERVE-CLI-BENCH` | Serve and latency/throughput/serve benchmark modes | T0 | `vllm/entrypoints/cli/serve.py:44`; `vllm/entrypoints/cli/benchmark/main.py:29` | separate binaries + explicit scheduler-capacity flags `examples/server/main.cpp:63,96,116,170`; `examples/bench/main.cpp:40,109`; `examples/bench/bench_core.h:96,468` | server help contract `examples/CMakeLists.txt:34`; benchmark `tests/examples/test_bench.cpp:15,48` | `planned: specs/cli-serve-bench.md` | `PARTIAL` | - |
160-
| `SERVE-GATE-ONLINE` | Same-corpus online correctness, TTFT/TPOT/ITL, throughput and peak-memory gate vs vLLM v0.25.0 | T0 | `vllm/benchmarks/serve.py:1,581-615`; [v0.25 audit](sync/2026-07-12-702f481.md); `tests/benchmarks/test_serve_cli.py:1` | Existing schema-v5 harness plus [trace controller](../include/vt/cuda/cuda_profiler_control.h#L13), [driver](../scripts/dgx-online-serving.sh#L14), batch-keyed [validator](../tools/bench/online_gate.py#L99), and fail-closed [c2 finalizer](../tools/bench/finalize_low_batch_trace.py#L1) | Immutable `3f256ab` remains **55/124**. Finalized `179a0fc` status `9e0143fa…7b57` proves exact before-state topology and matching FP4 tactics. Clean pushed `581d335` closes W1 BA F32-output core correctness/safety: merged/split 27B are byte-identical at **235/235 + 16/16**, strict memcheck is clean, and native 35B is **315/315**. BF16 output fails **233/235**; the c2 validator still needs merged/split contracts before exact 145-vs-193/component evidence. No speed credit. Host memory remains **22.920 GiB** CPU weights plus mmap pages; exact grid and 35B performance remain blocked | [online serving gate](specs/cuda-online-serving-gate.md); [merged GDN projections](specs/gdn-merged-input-projections.md) | `ACTIVE` | CLAIM-SERVE-GATE-1 |
162+
| `SERVE-GATE-ONLINE` | Same-corpus online correctness, TTFT/TPOT/ITL, throughput and peak-memory gate vs vLLM v0.25.0 | T0 | `vllm/benchmarks/serve.py:1,581-615`; [v0.25 audit](sync/2026-07-12-702f481.md); `tests/benchmarks/test_serve_cli.py:1` | Schema-v5 harness plus [trace controller](../include/vt/cuda/cuda_profiler_control.h#L13), one-lock mode-aware [driver](../scripts/dgx-online-serving.sh#L405), explicit [validator contracts](../tools/bench/online_gate.py#L113), and fail-closed [GDN BA finalizer](../tools/bench/finalize_gdn_ba_trace.py#L1) | Immutable `3f256ab` remains **55/124**. Finalized `179a0fc` proves the 193-vs-97 before-state and matching FP4 tactics; pushed `581d335` closes W1 BA core correctness/safety. The locally green harness preserves old 1,011-node evidence, requires merged **963/145 BF16** versus split **1,011/193**, records the toggle, pairs both arms with fresh vLLM under one lock, and accepts only an exact 48-BF16-only delta; **68/68** tool tests pass. Pushed-SHA GPU execution, BF16 output (233/235), component evidence, host-memory repair and exact grid remain open; no speed credit | [online serving gate](specs/cuda-online-serving-gate.md); [merged GDN projections](specs/gdn-merged-input-projections.md) | `ACTIVE` | CLAIM-SERVE-GATE-1 |
161163
| `SERVE-E2E-NIGHTLY` | Server conformance and real-model nightly suites for all release gates | T0 | `tests/entrypoints/openai/`; `tests/v1/e2e/`; `.buildkite/test-pipeline.yaml` | current unit/conformance tests only; no scheduled DGX suite | `tests/vllm/entrypoints/openai/test_conformance.cpp:1`; `tests/parity/test_qwen36_paged_engine.cpp:78`; `tests/parity/test_qwen27_paged_engine.cpp:110` | `planned: specs/server-e2e-nightly.md` | `INVENTORIED` | - |
162164
| `SERVE-CLI-CHAT` | Interactive chat and complete commands | T1 | `vllm/entrypoints/cli/main.py:18-34` has no direct chat/complete command at the pin; project extension | - | - | `planned: specs/cli-chat-complete.md` | `INVENTORIED` | - |
163165
| `SERVE-POOLING-ENDPOINTS` | Embeddings, pooling, score, rerank | T2 | `vllm/entrypoints/pooling/embed/api_router.py:25`; `vllm/entrypoints/pooling/scoring/api_router.py:1` | - | - | `planned: specs/pooling-endpoints.md` | `INVENTORIED` | - |

.agents/environment.md

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -87,8 +87,10 @@ inner 4096, state 128; context 262144.
8787
`~/work/vllm.cpp-gdn-ba/immutable-581d335…` passes the exact CUDA 13.0.88 /
8888
CUTLASS / Triton-AOT build, packed F32/BF16 capture/replay, strict memcheck,
8989
merged/split 27B and inert native-35B gates; BF16 projection output fails the token
90-
near-tie. Extend the exact-c2 harness, then close 145-vs-193 trace, rounding
91-
parity and the c2/c16 component before qkvz. Independently remove **22.920
90+
near-tie. The exact-c2 harness now carries explicit merged/split contracts,
91+
runs both paired arms under one lock and passes 68/68 local tool tests. Execute
92+
and finalize it from the pushed SHA, then close rounding parity and the c2/c16
93+
component before qkvz. Independently remove **22.920
9294
GiB** host-weight mirror and overlapping source pages. No 35B performance
9395
command runs before all 27B axes pass.
9496
- Keep the existing SGLang v0.5.13 P1 evidence immutable. The distinct

.agents/feature-matrix.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -108,7 +108,7 @@ and reference-engine performance.
108108

109109
| ID | Block | State | Grounded summary | Detailed evidence / spike |
110110
|---|---|---|---|---|
111-
| `QUANT-CUDA-GATES` | NVFP4 W4A16, NVFP4 W4A4, gate-specific FP8 W8A8 | `DONE` | support/correctness stays closed; performance remains `ACTIVE` at `3f256ab` **55/124**. Finalized `179a0fc` proves the FP4 tactic family matches; no quantization lever or speed credit follows. The separate [merged-GDN spike](specs/gdn-merged-input-projections.md) has BA-only W1 `GATING`: pushed `581d335` is F32-output core-correctness/safety green, BF16 output is not, and structure/performance remain pending | quant matrix §2 + [coverage spike](specs/quantization-coverage.md) |
111+
| `QUANT-CUDA-GATES` | NVFP4 W4A16, NVFP4 W4A4, gate-specific FP8 W8A8 | `DONE` | support/correctness stays closed; performance remains `ACTIVE` at `3f256ab` **55/124**. Finalized `179a0fc` proves the FP4 tactic family matches; no quantization lever or speed credit follows. The separate [merged-GDN spike](specs/gdn-merged-input-projections.md) has BA-only W1 `GATING`: pushed `581d335` is F32-output core-correctness/safety green and its explicit 145/193 c2 harness passes 68/68 locally, while BF16 output, immutable trace and component performance remain open | quant matrix §2 + [coverage spike](specs/quantization-coverage.md) |
112112
| `QUANT-GGUF` | llama.cpp encodings and output presets | `PARTIAL` | F32/Q4_0/Q8_0/Q3_K/Q4_K/Q5_K/Q6_K materialize; F16 was corrected to reader-only; CPU threadpool/chunked dispatch is correctness-gated but its B4 speed/RSS checkpoint is pending; no direct compute-in-quant or llama.cpp speed parity | quant matrix §1 |
113113
| `QUANT-VLLM-BREADTH` | generic FP8/MX, AWQ/GPTQ, CT integer, vendor methods, KV | `PARTIAL` | gate-specific implementations exist; generic dispatch/modes remain inventoried | quant matrix §§2-3 |
114114
| `QUANT-MLX` | affine Q2-8, MXFP4/MXFP8/NVFP4, QQ, mixed recipes/imports | `INVENTORIED` | required for Apple backend; no MLX runtime yet | quant matrix §4 |

0 commit comments

Comments
 (0)