You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: .agents/coordination.md
+2-2Lines changed: 2 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -113,7 +113,7 @@ time owns the GB10. Results without the lock for their entire run are discarded.
113
113
| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
114
114
|---|---|---|---|---|---|---|---|
115
115
| `CLAIM-PR3` | `KERNEL-GDN-AOT-BF16`, `KERNEL-GDN-SCRATCH` | root takeover of stopped `validate_pr3` / `complete_pr3` stream | primary recovery tree `/home/mudler/_git/vllm.cpp-pr3-validate`; 27B default/component integration in `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; DGX evidence `~/work/vllm.cpp-noPy` plus `~/work/vllm.cpp-nvfp4-small-m/debug/gdn-out-bf16-c16-ab-20260712` | recovery branch integrated into `main` by `a767188`; current checkpoint on `codex/nvfp4-small-m` | PR #3 files and rows `KERNEL-GDN-AOT-BF16`+`KERNEL-GDN-SCRATCH`; 27B-only `GdnOutDType` default/f32 override in `qwen3_5.cpp`; ledger/inventory/roadmap evidence. No 35B default change. GPU lock: residual trace/pool classification remains separate from FP4 W3 | `ACTIVE` | 2026-07-13 (vendored BF16 H32/H48 AOT/safety evidence and native 16/16 correctness remain green. Binding `3f256ab` has c16 total at 1.027889× but mean TPOT/ITL at 0.987450×. The dynamic scan ranks packed pure-decode GDN fusion after W3-C removes tactic-selection confounding. All 35B paths keep f32) |
116
-
| `CLAIM-SERVE-GATE-1` | `SERVE-GATE-ONLINE`, `KERNEL-GEMM-NVFP4-W4A4`, `KERNEL-ATTN-FA2` | root takeover | binding grid `~/work/vllm.cpp-online-gate/evidence/3f256abdbb558e162bf8a2196284deb119648560`; active worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; finalized c2 root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; W1 immutable root `~/work/vllm.cpp-gdn-ba/immutable-581d335fec2e5a96d9ccbb38c1ec001c39ac1789` | current checkpoint `codex/nvfp4-small-m` | Online-gate execution and harness work, including immutable evidence for the now-unclaimed `GATING` W1 BA implementation. qkvz code, loader-memory repair, exact grid, and 35B performance are excluded. Any A/B or trace series uses one `flock`; further `KERNEL-GEMM-BF16` production code requires a fresh `ACTIVE` claim | `ACTIVE` | 2026-07-14 (W1 implementation claim released at `GATING`. Clean pushed `581d335` closes F32-output core correctness/safety: merged/split 27B **235/235 + 16/16**, strict memcheck **590/590**, native 35B **315/315**. BF16 output fails **233/235**; exact-c2 merged/split harness, trace/component and rounding repair remain pending; binding 55/124, no speed credit) |
116
+
| `CLAIM-SERVE-GATE-1` | `SERVE-GATE-ONLINE`, `KERNEL-GEMM-NVFP4-W4A4`, `KERNEL-ATTN-FA2` | root takeover | binding grid `~/work/vllm.cpp-online-gate/evidence/3f256abdbb558e162bf8a2196284deb119648560`; active worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; finalized c2 root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; W1 safety root `~/work/vllm.cpp-gdn-ba/immutable-581d335fec2e5a96d9ccbb38c1ec001c39ac1789`; next trace root `~/work/vllm.cpp-gdn-ba-trace/<pushed-sha>` | current checkpoint `codex/nvfp4-small-m` | Online-gate execution and harness work, including immutable evidence for the now-unclaimed `GATING` W1 BA implementation. qkvz code, loader-memory repair, exact grid, and 35B performance are excluded. Any A/B or trace series uses one `flock`; further `KERNEL-GEMM-BF16` production code requires a fresh `ACTIVE` claim | `ACTIVE` | 2026-07-14 (W1 core correctness/safety remains closed at `581d335`; BF16 output fails **233/235**. Explicit merged/split c2 contracts, one-lock paired driver, provenance validation and completion-marker-last finalizer are locally green at **68/68** tool tests. Run/finalize the pushed SHA to prove 963/145 versus 1,011/193 before component work; binding 55/124, no speed credit) |
117
117
|`CLAIM-NVFP4-SMALL-M-2`|`KERNEL-GEMM-NVFP4-W4A4` (`W2`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; immutable source/evidence `~/work/vllm.cpp-nvfp4-small-m/b5c6e4fd65cdacea8f378e18ae101ebf521e8f01/w2` plus exact online gate |`codex/nvfp4-small-m`| W2 only: 32 tactics, merged CT gate/up and one-input `SiluAndMulFp4Quant`, with independent fallbacks and full correctness/safety/component/oracle gates |`RELEASED`| 2026-07-12 (implementation/correctness complete; strict old-oracle acceptance failed. The vLLM 0.24.0 ratios are historical diagnostics; W3 remains the trace-driven repair track under the new v0.25.0 denominator) |
118
118
|`CLAIM-NVFP4-SMALL-M-3`|`KERNEL-GEMM-NVFP4-W4A4` (`W3-C`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; oracle cache fixture under `tests/fixtures/nvfp4_flashinfer_v025_gb10/`; immutable C3/C3R/corrected component `~/work/vllm.cpp-nvfp4-persistent/d211b8f80fff831a712f0bfafa4f65f1abe1892d/evidence`|`codex/nvfp4-small-m`| W3-C document/import/atomicity, ready-map/lifecycle/5,000-us wiring and corrected same-plan gates only; other levers excluded |`RELEASED`| 2026-07-13 (W3-C reproduction control complete: six-process and corrected 12-leg components use identical 64/64 maps with zero tuning/misses. W3-E strict-fails 39/40 timing + 1/8 memory; no exact grid/35B performance) |
119
119
|`CLAIM-NVFP4-SMALL-M-4`|`KERNEL-GEMM-NVFP4-W4A4` (`W3-F`) | root | isolated worktree `/home/mudler/_git/vllm.cpp-nvfp4-small-m`; immutable source/evidence `~/work/vllm.cpp-nvfp4-device-alpha/7517af4f983fe322ac88ce2d9869e1441b7be3fd/`|`codex/nvfp4-small-m`| W3-F only: model-owned device F32 alpha, tensor op ABI/validation, `VT_FP4_DEVICE_ALPHA=0` fallback, ported tests, safety/model/node-trace and frozen-plan c2/c16 gates per [spike](specs/nvfp4-device-alpha.md). Quant/GDN/attention/host-weight changes excluded |`RELEASED`| 2026-07-13 (local/CUDA/operator/memcheck/model/trace gates pass. Completed 12-leg/612-request component is c2/c16 1.001967×/1.000144× but strict-fails 27/40 timing + 3/8 memory. No speed credit/exact grid/35B performance; the completed scan moves W3-G FA2 decode under CLAIM-SERVE-GATE-1) |
@@ -132,7 +132,7 @@ open. qkvz, exact-grid and 35B work stay excluded until those gates close.
132
132
133
133
| Priority | Row/block | Dependency | Next handoff | State |
134
134
|---|---|---|---|---|
135
-
| 1 |`KERNEL-GEMM-BF16` merged GDN projections for `SERVE-GATE-ONLINE`|`3f256ab` remains 55/124; pushed `581d335` closes F32-output core correctness/safety, BF16 output fails, and no trace/component speed credit exists | extend exact-c2 validation/driver for default and split; prove 145-vs-193 graph structure, resolve BF16 rounding, then complete c2/c16 before qkvz |`GATING`|
135
+
| 1 |`KERNEL-GEMM-BF16` merged GDN projections for `SERVE-GATE-ONLINE`|`3f256ab` remains 55/124; pushed `581d335` closes F32-output core correctness/safety, BF16 output fails, and the mode-aware trace harness is locally green with no GPU/speed credit | run/finalize its pushed-SHA merged/split series to prove 145-vs-193 graph structure, resolve BF16 rounding, then complete c2/c16 before qkvz |`GATING`|
136
136
| 2 |`SERVE-ASYNC-LLM` block (`ENG-CORE-BUSY-LOOP`, `SERVE-ASYNC-LLM`, `ENG-ASYNC-SCHED`, `ENG-PRIORITY-SCHED`) | joint spike accepted ([async-serving.md](specs/async-serving.md)); W1/W2/W4 implemented/`GATING`, W3 `READY`; complete control shows W3 is neutral for order-0 speed | retain W3 as unclaimed parity work until order 0 closes; do not fold it into the active kernel claim |`GATING`|
137
137
| 3 |`SERVE-E2E-NIGHTLY`|`SERVE-GATE-ONLINE` evidence where benchmarks overlap | write spike and CI/nightly split |`INVENTORIED`|
138
138
| 4 | C1 kernel drop-in alignment | accepted kernel-family inventory + [drop-in ABI spike](specs/dropin-kernel-abi.md)|`BACKEND-ABI-VT` W0 is implemented and CPU-green; run cross-build + flock-held GB10 runtime/capture/model/trace handoff, close scalar-forwarder/backend-shim debts, then migrate one family per independently gated checkpoint |`GATING`|
Copy file name to clipboardExpand all lines: .agents/engine-matrix.md
+5-3Lines changed: 5 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -20,8 +20,10 @@ vLLM qkvz+ba across 48 GDN layers. The accepted
20
20
[merged-projection spike](specs/gdn-merged-input-projections.md) makes
21
21
`KERNEL-GEMM-BF16` W1 `GATING`. Clean pushed `581d335` closes the one-owner
22
22
F32-output BA core correctness/safety gate; BF16 output fails the token near-tie,
23
-
while exact trace/component evidence remains pending. qkvz stays excluded.
24
-
No trace duration earns speed credit. Host PSS/RSS separately retains a
23
+
while exact trace/component evidence remains pending. The explicit mode-aware
24
+
c2 driver/contracts/finalizer pass 68/68 local tool tests; immutable GPU
25
+
execution is next. qkvz stays excluded. No trace duration earns speed credit.
26
+
Host PSS/RSS separately retains a
25
27
**22.920 GiB** CPU weight mirror plus source mmap residency.
26
28
27
29
| Area | Rows |`ANCHOR-BACKFILL`|`PARTIAL`|`SPIKE`|`READY`|`ACTIVE`|`GATING`|`INVENTORIED`|
@@ -157,7 +159,7 @@ claims it.
157
159
|`SERVE-C-ABI`| Stable LocalAI-style C FFI (17 exported symbols; blocking and nonblocking request handles) | T0 | Original project ABI; pinned vLLM has no C ABI |`include/vllm.h:143,181,207`; `src/capi/vllm_c.cpp:229,264,327,391`|`tests/capi/test_capi.cpp:320,428,505,574,606,640`; `tests/capi/test_dlopen.cpp:77,86`; `tests/capi/c_header_compile.c:1`|`planned: specs/c-api-library.md`|`ANCHOR-BACKFILL`| - |
158
160
|`SERVE-CPP-API`| Rich `LLM` and `AsyncLLM` C++ API | T1 |`vllm/entrypoints/llm.py:66,422`; `vllm/v1/engine/async_llm.py:70`| - | - |`planned: specs/cpp-api.md`|`INVENTORIED`| - |
159
161
|`SERVE-CLI-BENCH`| Serve and latency/throughput/serve benchmark modes | T0 |`vllm/entrypoints/cli/serve.py:44`; `vllm/entrypoints/cli/benchmark/main.py:29`| separate binaries + explicit scheduler-capacity flags `examples/server/main.cpp:63,96,116,170`; `examples/bench/main.cpp:40,109`; `examples/bench/bench_core.h:96,468`| server help contract `examples/CMakeLists.txt:34`; benchmark `tests/examples/test_bench.cpp:15,48`|`planned: specs/cli-serve-bench.md`|`PARTIAL`| - |
160
-
| `SERVE-GATE-ONLINE` | Same-corpus online correctness, TTFT/TPOT/ITL, throughput and peak-memory gate vs vLLM v0.25.0 | T0 | `vllm/benchmarks/serve.py:1,581-615`; [v0.25 audit](sync/2026-07-12-702f481.md); `tests/benchmarks/test_serve_cli.py:1` | Existing schema-v5 harness plus [trace controller](../include/vt/cuda/cuda_profiler_control.h#L13), [driver](../scripts/dgx-online-serving.sh#L14), batch-keyed [validator](../tools/bench/online_gate.py#L99), and fail-closed [c2 finalizer](../tools/bench/finalize_low_batch_trace.py#L1) | Immutable `3f256ab` remains **55/124**. Finalized `179a0fc` status `9e0143fa…7b57` proves exact before-state topology and matching FP4 tactics. Clean pushed `581d335` closes W1 BA F32-output core correctness/safety: merged/split 27B are byte-identical at **235/235 + 16/16**, strict memcheck is clean, and native 35B is **315/315**. BF16 output fails **233/235**; the c2 validator still needs merged/split contracts before exact 145-vs-193/component evidence. No speed credit. Host memory remains **22.920 GiB** CPU weights plus mmap pages; exact grid and 35B performance remain blocked | [online serving gate](specs/cuda-online-serving-gate.md); [merged GDN projections](specs/gdn-merged-input-projections.md) | `ACTIVE` | CLAIM-SERVE-GATE-1 |
162
+
| `SERVE-GATE-ONLINE` | Same-corpus online correctness, TTFT/TPOT/ITL, throughput and peak-memory gate vs vLLM v0.25.0 | T0 | `vllm/benchmarks/serve.py:1,581-615`; [v0.25 audit](sync/2026-07-12-702f481.md); `tests/benchmarks/test_serve_cli.py:1` | Schema-v5 harness plus [trace controller](../include/vt/cuda/cuda_profiler_control.h#L13), one-lock mode-aware [driver](../scripts/dgx-online-serving.sh#L405), explicit [validator contracts](../tools/bench/online_gate.py#L113), and fail-closed [GDN BA finalizer](../tools/bench/finalize_gdn_ba_trace.py#L1) | Immutable `3f256ab` remains **55/124**. Finalized `179a0fc` proves the 193-vs-97 before-state and matching FP4 tactics; pushed `581d335` closes W1 BA core correctness/safety. The locally green harness preserves old 1,011-node evidence, requires merged **963/145 BF16** versus split **1,011/193**, records the toggle, pairs both arms with fresh vLLM under one lock, and accepts only an exact 48-BF16-only delta; **68/68** tool tests pass. Pushed-SHA GPU execution, BF16 output (233/235), component evidence, host-memory repair and exact grid remain open; no speed credit | [online serving gate](specs/cuda-online-serving-gate.md); [merged GDN projections](specs/gdn-merged-input-projections.md) | `ACTIVE` | CLAIM-SERVE-GATE-1 |
161
163
|`SERVE-E2E-NIGHTLY`| Server conformance and real-model nightly suites for all release gates | T0 |`tests/entrypoints/openai/`; `tests/v1/e2e/`; `.buildkite/test-pipeline.yaml`| current unit/conformance tests only; no scheduled DGX suite |`tests/vllm/entrypoints/openai/test_conformance.cpp:1`; `tests/parity/test_qwen36_paged_engine.cpp:78`; `tests/parity/test_qwen27_paged_engine.cpp:110`|`planned: specs/server-e2e-nightly.md`|`INVENTORIED`| - |
162
164
|`SERVE-CLI-CHAT`| Interactive chat and complete commands | T1 |`vllm/entrypoints/cli/main.py:18-34` has no direct chat/complete command at the pin; project extension | - | - |`planned: specs/cli-chat-complete.md`|`INVENTORIED`| - |
Copy file name to clipboardExpand all lines: .agents/feature-matrix.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -108,7 +108,7 @@ and reference-engine performance.
108
108
109
109
| ID | Block | State | Grounded summary | Detailed evidence / spike |
110
110
|---|---|---|---|---|
111
-
|`QUANT-CUDA-GATES`| NVFP4 W4A16, NVFP4 W4A4, gate-specific FP8 W8A8 |`DONE`| support/correctness stays closed; performance remains `ACTIVE` at `3f256ab`**55/124**. Finalized `179a0fc` proves the FP4 tactic family matches; no quantization lever or speed credit follows. The separate [merged-GDN spike](specs/gdn-merged-input-projections.md) has BA-only W1 `GATING`: pushed `581d335` is F32-output core-correctness/safety green, BF16 output is not, and structure/performance remain pending| quant matrix §2 + [coverage spike](specs/quantization-coverage.md)|
111
+
|`QUANT-CUDA-GATES`| NVFP4 W4A16, NVFP4 W4A4, gate-specific FP8 W8A8 |`DONE`| support/correctness stays closed; performance remains `ACTIVE` at `3f256ab`**55/124**. Finalized `179a0fc` proves the FP4 tactic family matches; no quantization lever or speed credit follows. The separate [merged-GDN spike](specs/gdn-merged-input-projections.md) has BA-only W1 `GATING`: pushed `581d335` is F32-output core-correctness/safety green and its explicit 145/193 c2 harness passes 68/68 locally, while BF16 output, immutable trace and component performance remain open| quant matrix §2 + [coverage spike](specs/quantization-coverage.md)|
112
112
|`QUANT-GGUF`| llama.cpp encodings and output presets |`PARTIAL`| F32/Q4_0/Q8_0/Q3_K/Q4_K/Q5_K/Q6_K materialize; F16 was corrected to reader-only; CPU threadpool/chunked dispatch is correctness-gated but its B4 speed/RSS checkpoint is pending; no direct compute-in-quant or llama.cpp speed parity | quant matrix §1 |
0 commit comments