Skip to content

test(rocm): ReshapeAndCache->PagedAttention composition at real dims (issue #41) - #497

Draft
VikashLoomba wants to merge 1 commit into
mudler:mainfrom
VikashLoomba:row/ROCM-ATTN-COMPOSE-TEST
Draft

test(rocm): ReshapeAndCache->PagedAttention composition at real dims (issue #41)#497
VikashLoomba wants to merge 1 commit into
mudler:mainfrom
VikashLoomba:row/ROCM-ATTN-COMPOSE-TEST

Conversation

@VikashLoomba

Copy link
Copy Markdown
Contributor

Row

BACKEND-ROCM — test-hardening only (one additive cross-device case). Issue #41.

What changed

Adds a cross-device case for the KV-cache composition the in-tree suite doesn't cover: the existing paged-attention case hand-builds a contiguous cache, but the real model path writes KV via ReshapeAndCache and reads it back via PagedAttention. This case is that composition at real Qwen3.5-0.8B dims (Dh=256, Hq=8, Hkv=2, block_size 16), a shuffled block table, and a non-sequential slot mapping — the layout a stride/scatter bug would live in and the contiguous case cannot see.

Surfaced by the #41 Qwen3.5-0.8B divergence investigation: with this composition passing, every piece of the ROCm attention path validates in isolation, which is what localizes the residual divergence to bf16-softmax accumulation rather than a kernel defect (full causal chain in this #41 comment).

Evidence (4× gfx1100, ROCm 7.14, Release)

  • New case: 7/7 vs the CPU oracle (runs on ROCm; the composition exercised end to end)
  • Full cross-device suite green
  • agent-preflight.sh --staged green; check-commit-trailers green

Speed claims

  • This PR makes NO speed claim.

Honest gaps

  • Test-only; no behavior change. The case validated the path as-is — it did not surface a defect (that was the point: it closes a coverage gap the divergence investigation needed ruled out).

…udler#41)

The in-tree paged-attention case hand-builds a contiguous KV cache; the real
model path writes it with ReshapeAndCache and reads it back. This case is
that composition at real model dims (Dh=256, Hq=8, Hkv=2, block_size 16)
with a shuffled block table and non-sequential slot mapping — the layout a
stride/scatter bug would live in and the contiguous case cannot see.

Surfaced by the mudler#41 Qwen3.5-0.8B divergence investigation: every
compositional piece of the ROCm attention path now validates in isolation,
which is what localizes the residual divergence to bf16-softmax accumulation
rather than a kernel defect.

Evidence (4x gfx1100, ROCm 7.14, Release): the new case passes 7/7 vs the
CPU oracle; full cross-device suite green.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: pi:kimi-k3 [pi]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant