MSVC+CUDA llama.cpp fork for Windows: KDA/GDN, long-context KV placement, TurboQuant/TCQ when measured, speculative draft. Lab defaults — TriAttention opt-in only, not recommended.
-
Updated
Jul 28, 2026 - C++
MSVC+CUDA llama.cpp fork for Windows: KDA/GDN, long-context KV placement, TurboQuant/TCQ when measured, speculative draft. Lab defaults — TriAttention opt-in only, not recommended.
Pre-built PyTorch wheels and build scripts for NVIDIA DGX Spark (GB10, sm_121, Blackwell, CUDA 13.0, ARM64)
Build the GPU inference stack from source, repeatably: for Blackwell sm_120 on CUDA 13.x and Python 3.14. Machine-readable build state, generated patches, and an abductive-triage skill for build failures.
Hunyuan3D-2 fork — image→textured 3D→sliced STL + part segmentation. RTX 50-series (Blackwell/sm_120), CUDA 13.0, Python 3.12, PyTorch 2.11+cu130.
Windows NVIDIA-only Triton 3.7.0 build pipeline for RTX 5090 / Blackwell sm_120a, with FP8 tl.dot validation and peak benchmark results.
Accelerate Kimi Delta Attention computations with high-performance CUTLASS kernels designed for NVIDIA SM90 architectures and beyond.
Add a description, image, and links to the cuda-13 topic page so that developers can more easily learn about it.
To associate your repository with the cuda-13 topic, visit your repo's landing page and select "manage topics."