feat(comm): add selected-record PCIe exchange#2
Draft
FujitsuPolycom wants to merge 1 commit into
Draft
Conversation
Add a direct CUDA-IPC exchange for destination-selected fixed-width records, expose it through the Sparkinfer PCIe facade, and preserve runtime-JIT CUDA sources in built wheels. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: FujitsuPolycom <87842395+FujitsuPolycom@users.noreply.github.com>
FujitsuPolycom
force-pushed
the
codex/sparse-ckv-p2p-transport
branch
from
July 21, 2026 04:24
a7ebc4b to
f3d5c14
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The public API is exposed from
sparkinfer.comm.pcieasSelectedRecordExchange. GLM/vLLM union, remap, prefetch, and attention policyremain outside Sparkinfer.
Sparkinfer rebase
The branch is based on
local-inference-lab/sparkinfer@ec2cd4d. Python and CUDAsources live under
sparkinfer/comm/pcie, use Sparkinfer-prefixed JIT and testgates, and are included in the built wheel. No
b12x.distributedcompatibilitypackage is added.
Validation
Test host: ai01 is an ASUS Pro WS WRX90E-SAGE SE with a Threadripper PRO
9965WX, 128 GiB RAM, and 4x RTX PRO 6000 Blackwell 96 GB GPUs at 400 W.
Scope
Transport only. The copy-engine variant is intentionally stacked in the next PR.
AI-assisted implementation under human direction; the diff and validation
results were reviewed before publication.
Design document
GLM-5.2 Sparse CKV under Decode Context Parallelism is the canonical architecture, transport-contract, validation, and upstream-decomposition reference for this PR stack.