-
Notifications
You must be signed in to change notification settings - Fork 224
OpenJul 24, 2026
No due date
•Last updated 61% complete
List view
0 of 19 selected 0 issues of 19 selected
Tracking: cuDNN Python SDPA API Integration
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#309 In NVIDIA/cudnn-frontend;Tracking: Have cutedsl MoE kernels perf dashboard
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.Status: Open.#320 In NVIDIA/cudnn-frontend;Missing THD support for cuDNN 9.23 + cuDNN FE 1.24 for D=256 bwd on sm10.x
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-backendcuDNN backend API, graph execution, descriptors, engines, or backend integration.cuDNN backend API, graph execution, descriptors, engines, or backend integration.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#276 In NVIDIA/cudnn-frontend;cuDNN frontend grouped GEMM fusion API is restricted to PyTorch
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#260 In NVIDIA/cudnn-frontend;[PyTorch][CUDNN][Windows] Version 1.22.1 fails to compile with NV_CUDNN_FRONTEND_USE_DYNAMIC_LOADING
mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.open-todoOpen work that is ready to be picked up, scheduled, or prioritized.Open work that is ready to be picked up, scheduled, or prioritized.Status: Open.#261 In NVIDIA/cudnn-frontend;experimental: grouped gemm AOT
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Draft (not ready).[Feature request] Include CMake config into the package
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#340 In NVIDIA/cudnn-frontend;Kernel dies executing sample 26 (layernorm + ReLU bitmask) after cudnn.pygraph IR rewrite (fd2cdd18)
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.open-verify-to-closeOpen work with a proposed fix that needs verification before closing.Open work with a proposed fix that needs verification before closing.orig-qaReported or owned by quality assurance, validation, or testing.Reported or owned by quality assurance, validation, or testing.Status: Open.#361 In NVIDIA/cudnn-frontend;[feature request] Add blackwell HSTU mha kernel.
mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#369 In NVIDIA/cudnn-frontend;[Bug]: cutedsl kernel fails on upgrade to 4.6.0
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.Support installing python bindings via cmake --install (fixes #149)
cat-ciCI failures, test flakiness, workflow breakage, or automation issues.CI failures, test flakiness, workflow breakage, or automation issues.cat-infraBuild, packaging, tooling, dependency, release, or repository maintenance work.Build, packaging, tooling, dependency, release, or repository maintenance work.mod-infraInfrastructure, CI/CD, build systems, packaging, releases, or repo maintenance.Infrastructure, CI/CD, build systems, packaging, releases, or repo maintenance.orig-externalReported or requested by an external user, customer, or community contributor.Reported or requested by an external user, customer, or community contributor.Status: Open (in progress).[CI][DSA][Blackwell] oss:rel shard1 sparse attention backward Tensor-likes mismatch
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-qaReported or owned by quality assurance, validation, or testing.Reported or owned by quality assurance, validation, or testing.Status: Open.#418 In NVIDIA/cudnn-frontend;DSA: fix SM100 sink normalization for 576/512 backward
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 421#421 In NVIDIA/cudnn-frontend;feat: add BF16 grouped GEMM MoE kernels
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.Status: Open (in progress).Support F64 compute data type for convolutions
mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-externalReported or requested by an external user, customer, or community contributor.Reported or requested by an external user, customer, or community contributor.Status: Open (in progress).test_mhas_v2: extend bwd random head-dim coverage to d=256
mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).Fix stream-ordering and input-validation holes in the SM100 DSA backward interface
mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-externalReported or requested by an external user, customer, or community contributor.Reported or requested by an external user, customer, or community contributor.Status: Open (in progress).CSA: add fused Compressor forward+backward CuTe-DSL kernels (ported from Megatron-LM)
mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-externalReported or requested by an external user, customer, or community contributor.Reported or requested by an external user, customer, or community contributor.Status: Open (in progress).DSA backward SM100: reject topk_length entries < 1 (a zero entry hangs the kernel)
mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-externalReported or requested by an external user, customer, or community contributor.Reported or requested by an external user, customer, or community contributor.Status: Open (in progress).