Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
57 commits
Select commit Hold shift + click to select a range
65e79ee
feat(ltx-2.5): port the prompt-side AdaLN, and stop the opt-in from c…
localai-bot Aug 13, 2026
67e53e7
record(TOOLS-PARSER-BREADTH): backfill the 41 parsers already shipped…
localai-bot Aug 13, 2026
48412ae
spec(ENG-MM-INPUT-PIPELINE): map the mm CONFIG seam, including the RE…
localai-bot Aug 13, 2026
456e32c
merge: origin/main into row/LTX25-IMAGE-COND (#644)
mudler Aug 13, 2026
11cc1d5
fix(main-red): route LTX-2.5's device question through the platform s…
localai-bot Aug 13, 2026
eb58fd5
fix(ltx-2.5): the keyframe refusal named a FALSE reason, and a test p…
mudler Aug 13, 2026
dd41622
docs: align quantization and registry counts (#575)
localai-org-maint-bot Aug 13, 2026
7ba9a67
feat(rocm): implement the vt::Backend graph-capture seam on hipGraph …
joral Aug 13, 2026
7965f12
fix(GATE-PR-SIZE-BINARY): retire the fail-closed binary guard — a gol…
localai-bot Aug 13, 2026
782264c
spec(ENG-UPSTREAM-OMNI-PIN): pin vLLM-Omni onto the oracle registry (…
localai-bot Aug 13, 2026
b338e93
spec(MODEL-MM-indextts2): scope IndexTTS-2.5, the first audio-generat…
localai-bot Aug 13, 2026
43a6c55
W2 (KERNEL-SSM-MAMBA): the Mamba2 SSD CUDA arm — 9/9 mutations caught…
localai-bot Aug 13, 2026
4d0c399
W2 (MODEL-NEMOTRON-H): the non-gated relu² MoE expert — one GEMM, not…
localai-bot Aug 13, 2026
73d217d
feat(MODEL-MM-indextts2): W1 vocoder core, W2 GPT-2 backbone, W6a spe…
localai-bot Aug 13, 2026
f91d543
feat(MODEL-MM-indextts2): W3 first — CAMPPlus primitives gated on ups…
localai-bot Aug 13, 2026
e6ede74
feat(MODEL-MM-indextts2): W3 — CAMPPlus layer stack, and a defect the…
localai-bot Aug 13, 2026
a3aa02e
spec(MODEL-MUSIC-MUSIC3): scope MiniMax-Music3, the first row whose o…
localai-bot Aug 13, 2026
1e2a0af
feat(MODEL-MM-indextts2): W3 — FCM 2-D front end, verified on x86_64 …
localai-bot Aug 14, 2026
7138d16
test(#669): pin that the token budget SPLITS a prefill wave, and labe…
localai-bot Aug 14, 2026
87dbe37
feat(MODEL-MM-indextts2): CAMPPlus COMPLETE, plus a gate hole mutatio…
localai-bot Aug 14, 2026
5e010c9
feat(MODEL-MM-indextts2): w2v-bert-2.0 Conformer FFN + conv module (#…
localai-bot Aug 14, 2026
8534b88
feat(ENG-MM-INPUT-PIPELINE): multimodal input limits and the refusal …
localai-bot Aug 14, 2026
15e0311
spec(MODEL-MUSIC-MUSIC3): Music3 is a SpeechEngine family, not a new …
localai-bot Aug 14, 2026
487263b
feat(MODEL-MM-indextts2): w2v-bert relative-key self-attention (#634)…
localai-bot Aug 14, 2026
c57ca36
feat(MODEL-MM-indextts2): assemble the w2v-bert Conformer encoder lay…
localai-bot Aug 14, 2026
5562c47
feat(MODEL-MM-indextts2): w2v-bert semantic front end COMPLETE (#634)…
localai-bot Aug 14, 2026
e54a242
feat(TOOLS-PARSER-BREADTH): register `inkling` — the ported engine ha…
localai-bot Aug 14, 2026
bf33b75
feat(MODEL-MM-indextts2): EnhancedCodec's factorized VQ — the semanti…
localai-bot Aug 14, 2026
ba0732d
feat(MODEL-MM-indextts2): VocosBackbone — EnhancedCodec COMPLETE (#63…
localai-bot Aug 14, 2026
bbdccee
feat(MODEL-MM-indextts2): S2Mel length-regulator primitives (#634) (#…
localai-bot Aug 14, 2026
3d8e3f9
feat(MODEL-MM-indextts2): S2Mel flow-matching scaffolding, and a floa…
localai-bot Aug 14, 2026
34dc578
oracle(MODEL-MUSIC-MUSIC3): the diffusers oracle GENERATES AUDIO, and…
localai-bot Aug 14, 2026
8d0c277
feat(MODEL-MUSIC-MUSIC3): W1 — the modular checkpoint loader, and a d…
localai-bot Aug 14, 2026
3573428
feat(MODEL-MM-indextts2): adaLN conditioning and the DiT FinalLayer (…
localai-bot Aug 14, 2026
ce8c8bf
W4 (MODEL-NEMOTRON-H): the hybrid forward — and the inherited gate th…
localai-bot Aug 14, 2026
373aa12
spec(BACKEND-ROCM): the ROCm head_dim=128 decode arm, the ROCm half o…
joral Aug 14, 2026
a431364
feat(MODEL-MM-indextts2): S2Mel DiT block primitives — RMSNorm, adaLN…
localai-bot Aug 14, 2026
a23ba5b
feat(MODEL-MM-indextts2): assemble the S2Mel DiT block (#634) (#722)
localai-bot Aug 14, 2026
60d2574
feat(MODEL-MM-indextts2): talker embedding scaffolding — every stage …
localai-bot Aug 14, 2026
fd83efb
feat(MODEL-MM-indextts2): compose the six stages, and a transposed la…
localai-bot Aug 14, 2026
7a0e6c8
record(MODEL-MUSIC-MUSIC3): the row is ACTIVE, and the spec now meets…
localai-bot Aug 14, 2026
22367c5
feat(MODEL-MM-indextts2): the REAL shape contract, and four findings …
localai-bot Aug 14, 2026
3d64a3d
record(mm-serving): multiple images are TRUNCATED, not "not routed" (…
localai-bot Aug 14, 2026
8e463e7
fix(MODEL-MM-indextts2): the emotion model needs NO port, and the gua…
localai-bot Aug 14, 2026
e1e2de9
record(MODEL-MM-indextts2): 1712 tensors read from headers, and the s…
localai-bot Aug 14, 2026
5498b4a
Qwen3.5/3.8 text-only arms: register Qwen3_5[Moe]ForCausalLM, resolve…
localai-bot Aug 14, 2026
67dc0aa
feat(MODEL-MM-indextts2): port the DiT's WaveNet FINAL LAYER, the sta…
localai-bot Aug 14, 2026
33c18f9
feat(MODEL-MM-indextts2): the DiT tail composes end to end, and the w…
localai-bot Aug 14, 2026
d0bfea5
feat(MODEL-MM-indextts2): the DiT's U-Net skip routing, RECORDED from…
localai-bot Aug 14, 2026
f8cbc23
fix(#664): the video registry's existence probes stop reaching Window…
localai-bot Aug 14, 2026
0011bed
fix(#720): M_PI is not defined by MSVC — four TUs move to std::number…
localai-bot Aug 14, 2026
25861b6
feat(MODEL-MM-indextts2): convert the checkpoints OFFLINE, and refuse…
localai-bot Aug 14, 2026
b203871
spec(MODEL-MM-dots3-note): scope dots3-note, a DeepSeek-V3.2 text tow…
localai-bot Aug 14, 2026
5e646d9
feat(MODEL-MM-indextts2): the ported DiT runs on the REAL shipped wei…
localai-bot Aug 14, 2026
d374e83
feat(MODEL-MM-indextts2): the DiT conditioning front end, and a tenso…
localai-bot Aug 14, 2026
32d82c6
feat(MODEL-MM-indextts2): the DiT transformer stack, so the S2Mel est…
localai-bot Aug 14, 2026
56c2309
merge: main @ 32d82c64d into row/LTX25-IMAGE-COND-FIX (#644)
mudler Aug 14, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .agents/claims/CLAIM-MODEL-DOTS3-NOTE-W0.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# CLAIM-MODEL-DOTS3-NOTE-W0

| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
|---|---|---|---|---|---|---|---|
| `CLAIM-MODEL-DOTS3-NOTE-W0` | `MODEL-MM-dots3-note-dots3-note-for-causal-lm` (`SPIKE`) | Claude Code (opus-5), operator role | isolated worktree `/home/mudler/_git/wt-dots3-note`; oracle checkout read-only at `${VLLM_SOURCE}` = `/home/mudler/_git/vllm` (fetched to `origin/main` for the beyond-pin read, pin itself untouched) | `row/MODEL-MM-dots3-note-dots3-note-for-causal-lm`, issue [#699](https://github.com/mudler/vllm.cpp/issues/699) | Owns: NEW `.agents/specs/dots3-note.md`; the two dots3 rows in `.agents/model-matrix.md` (the `MODEL-MM` row at `SPIKE` and the `MODEL-SPEC-dots3-note-dots3-note-mtp` row left `INVENTORIED` and therefore not listed as a claimed row above, the engaged-architecture checklist entry, the rollup counts `373 -> 375` and the beyond-pin ratchet `369 -> 371`); the matching `MODEL` ratchet `373 -> 375` in `scripts/check-agent-record.py`; the `#699` row in `.agents/roadmap_v1.md`; the dots3 lines in `docs/FEATURES.md`. EXCLUDES: all engine code — no `src/`, `include/` or `tests/` change is in W0's scope, and W1 onward are dispatched to fresh implementers from the committed spec rather than written here. Also excludes any pin advance: this row is beyond-pin and says so, but advancing `555967922` is a sync cycle of its own and is not claimed here | `ACTIVE` | 2026-08-14 — developer-directed scope ("open the issue and write the spec, and merge it directly and add to roadmap. Use 192.168.68.23 as cuda host to run verifications e2e"), which records merge authority for W0 and designates Thor as the e2e CUDA host. W0 spec committed. **The row is BLOCKED on spec §6.4**, a developer decision between renting 8xH100, accepting unit-gated bricks with the e2e gate recorded as owed, or parking. W0.5 (provision Thor: no nvcc, no cmake, no ninja, no venv as probed 2026-08-14) is the first dispatchable task if option B is chosen. No W1 dispatch until then. |
5 changes: 5 additions & 0 deletions .agents/claims/CLAIM-MODEL-MUSIC3-W0.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# CLAIM-MODEL-MUSIC3-W0

| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
|---|---|---|---|---|---|---|---|
| `CLAIM-MODEL-MUSIC3-W0` | `MODEL-MUSIC-minimax-music3-mini-max-music3-for-conditional-generation` (`SPIKE`) | Claude Code (opus-5), operator role | isolated worktree `/home/mudler/_git/wt-music3`; oracle checkouts read-only at `/home/mudler/_git/sglang-omni` | `row/MODEL-MUSIC-MINIMAX-MUSIC3`, issue [#672](https://github.com/mudler/vllm.cpp/issues/672) | Owns: NEW `.agents/specs/minimax-music3.md`; the Music3 rows in `.agents/roadmap_v1.md` and `.agents/model-matrix.md` (detailed row, checklist entry, rollup counts, and the MODEL ratchet `370 -> 371`); NEW `.agents/oracles/sglang-omni.md` plus the `diffusers` pin advance to the PR #14456 head and the AGENTS.md registry row that admits `sglang-omni`. EXCLUDES: all engine code — no `src/`, `include/` or `tests/` change is in W0's scope, and W1 onward are dispatched to fresh implementers from the committed spec rather than written here. Also excludes the native `AbabForCausalLM` checkpoint layout, which the spec refuses by name and records as owed rather than silently mis-loading | `ACTIVE` | 2026-08-13 — **MERGE AUTHORITY RECORDED**: developer-directed "land minimax music 3 support complete, to vllm.cpp, wired to the ABI and to the example http server, merge to main, tested e2e". W0 spec committed. The oracle stand-up (proving `diffusers` builds and runs this model, which flips `.agents/oracles/diffusers.md` to `gateable = yes`) and the sample-rate resolution of spec §1.1 are the remaining W0 work; W1 is not dispatched until both land. |
5 changes: 5 additions & 0 deletions .agents/claims/CLAIM-ROCM-DECODE-ATTN-D128.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# CLAIM-ROCM-DECODE-ATTN-D128

| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
|---|---|---|---|---|---|---|---|
| `CLAIM-ROCM-DECODE-ATTN-D128` | `BACKEND-ROCM` (`ACTIVE`) | Claude Code (sonnet-5), helper role | worktree `rdna3-kernel-porting-b9ec47`, real gfx1200 hardware (AMD Radeon RX 9060 XT, RDNA4, 32 CU), `$GPU_LOCK` respected | `row/ROCM-DECODE-ATTN-D128-SPEC` (this spec; the implementation follows on `row/ROCM-DECODE-ATTN-D128-IMPL`, stacked), base `main` `fafa16f0`; issue [#382](https://github.com/mudler/vllm.cpp/issues/382) (the ROCm half; the CUDA half landed as [PR #425](https://github.com/mudler/vllm.cpp/pull/425), `66399617`), motivated by [#488](https://github.com/mudler/vllm.cpp/issues/488). NOTE: #382 is filed against the cross-backend kernel row (state `ANCHOR-BACKFILL`), while this claim's Row ID is the `ACTIVE` backend row whose code it edits — `check-agent-record` requires an active claim to name a `SPIKE`/`ACTIVE` row, so the two deliberately differ | Owns ONLY: the `LoadRowEplBf16`/`StoreRowEplBf16` `EPL=4` case, the `VT_ATTN_DECODE_D128` gate (default OFF, same flag/default/reason as the merged CUDA arm), the `bf16_decode_opt`/`decode_gqa` gate extensions and the two `d==128` launch-dispatch branches in `src/vt/rocm/rocm_paged_attn.hip`; the new "Qwen3 geometry (bf16, GQA 2, head_dim 128)" case in `tests/vt/test_backend_cross_device.cpp` and its two flag-on ctest registrations in `tests/CMakeLists.txt`; `.agents/specs/rocm-decode-attn-d128.md` and this claim file. **NON-COLLISION:** disjoint from `CLAIM-ROCM-SKINNY-GEMM-GFX1200` (different files: `rocm_skinny_gemm.hip`/`rocm_matmul_hipblaslt.hip` vs `rocm_paged_attn.hip`), not stacked on any other branch. EXCLUDED: **the flip to default-ON on either backend** (owes the near-tie razor + distributional gate + golden regen, and per the spec §5 cross-arch reversal must be argued per backend — this is what keeps #382 open), rocWMMA for `d=128` (separate claim, separate spec, separate issue), `qg=4`/`qg=8` GQA fusion at any `d` (pre-existing, board-independent gap), any `d=128` prefill path, and the 8 pre-existing unrelated `ctest` failures (`vt: no kernel for op 63 on device type 5`) | `ACTIVE` | 2026-08-12 — **reconciled against the existing record before landing**, per the re-verify-before-claiming rule: #382 already named this exact defect and PR #425 had already merged the CUDA half, so this became a mirror of merged work rather than new design, and was re-gated from default-ON to **default OFF behind `VT_ATTN_DECODE_D128`** — the merged arm's own flag, default and stated reason (warp-strided online softmax reduces the KV sequence in a different order, so a greedy anchor can move at an exact bf16 tie; OFF keeps every golden byte-identical). gfx1200-verified: `ctest -R 'rocm\|cross_device'` **6/6** including two new flag-on registrations (verified non-vacuous: 1 case, 6 assertions, not zero); full `ctest` 385/393 with the 8 failures independently confirmed pre-existing. Gate exercised **both directions on one binary** — Qwen3-0.6B @1024 ctx TPOT 44.82/44.82 ms OFF vs 12.80/12.60 ms ON = **3.53x**; decode throughput +42.7% / +25.0% / +17.8% on 0.6B / 1.7B / 4B. **Carried finding:** #382 measured this same `EPL=4` arm **1.6x slower** on sm_110 where we measure it 3.5x faster — recorded, not reconciled; it is why the default-ON flip must be argued per backend. Rebased from `bbc482a2` onto `main` `fafa16f0` (167 commits), which required reformatting `Assisted-by` for the `check-commit-trailers` gate that landed in between, and de-linking §7's forward reference to the rocWMMA spec — that spec now lands on its own branch, so a markdown link to it fails `check-agent-record` as a dangling link. Spec content otherwise byte-identical. Re-gated on the new base, gfx1200: build 783/783, `ctest -R 'rocm\|cross_device'` 6/6, the new case non-vacuous under both flags (1 case, 6 assertions), full `ctest` with 8 pre-existing `kSharedExpertGate` (`OpId(63)`) failures owed to unmerged PR #509. `agent-preflight` fails 11, set-identical to a clean `fafa16f0` baseline. Spec PR open; implementation PR follows. |
11 changes: 6 additions & 5 deletions .agents/engine-matrix.md

Large diffs are not rendered by default.

Loading
Loading