-
Notifications
You must be signed in to change notification settings - Fork 223
Pull requests: NVIDIA/cudnn-frontend
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Norm Samples updates for B,S,H style tensor inputs
#432
opened Jul 23, 2026 by
rmhaskarnvidia
Loading…
2 tasks done
Support mixed-form sequence lengths in SDPA forward (cuDNN 9.26+)
cat-enhancements
mod-frontend
cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.
orig-nv-eng
Reported or requested by NVIDIA engineering.
#430
opened Jul 23, 2026 by
egilliam-nv
Collaborator
Loading…
2 tasks done
Fix stream-ordering and input-validation holes in the SM100 DSA backward interface
#429
opened Jul 23, 2026 by
zkyue
Contributor
Loading…
2 tasks done
CSA: add fused Compressor forward+backward CuTe-DSL kernels (ported from Megatron-LM)
#427
opened Jul 23, 2026 by
zkyue
Contributor
Loading…
Fix a latent SMEM handoff race in the SM100 DSA indexer backward kernel
#426
opened Jul 23, 2026 by
zkyue
Contributor
Loading…
2 tasks done
test_mhas_v2: extend bwd random head-dim coverage to d=256
#425
opened Jul 22, 2026 by
vedaanta
Collaborator
Loading…
Support F64 compute data type for convolutions
cat-enhancements
mod-frontend
cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.
orig-external
Reported or requested by an external user, customer, or community contributor.
DSA: fix SM100 sink normalization for 576/512 backward
cat-bug
Reports of incorrect behavior, crashes, regressions, or unexpected results.
orig-nv-eng
Reported or requested by NVIDIA engineering.
DSA indexer forward: add a lean SM100 fast path for the head_dim=128 / qhead_per_kv_head=64 regime (up to 1.46x)
#416
opened Jul 21, 2026 by
zkyue
Contributor
Loading…
feat: add BF16 grouped GEMM MoE kernels
cat-feature
Requests for new functionality, APIs, examples, or behavior improvements.
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
Reject FP8/MXFP8 SDPA forward combinations exposed to a cuDNN 9.24 split-KV bug
#414
opened Jul 20, 2026 by
vedaanta
Collaborator
Loading…
Add version-adaptive nvvm.atomicrmw wrapper for cutlass-dsl 4.5.x compat
#411
opened Jul 20, 2026 by
Anerudhan
Collaborator
Loading…
Fix IMA caused by OOB lanes in varlen indexer top-k
cat-bug
Reports of incorrect behavior, crashes, regressions, or unexpected results.
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
orig-nv-eng
Reported or requested by NVIDIA engineering.
Use professional wording in env report docs
#409
opened Jul 20, 2026 by
liujane-dev
Collaborator
Loading…
2 tasks done
DSA: fix 576-wide d_sink, test/document top-k index semantics, B300 benchmarks
#405
opened Jul 17, 2026 by
vedaanta
Collaborator
Loading…
Support installing python bindings via cmake --install (fixes #149)
cat-ci
CI failures, test flakiness, workflow breakage, or automation issues.
cat-infra
Build, packaging, tooling, dependency, release, or repository maintenance work.
mod-infra
Infrastructure, CI/CD, build systems, packaging, releases, or repo maintenance.
orig-external
Reported or requested by an external user, customer, or community contributor.
Add FLOOR_MOD (floored modulo) pointwise mode
cat-feature
Requests for new functionality, APIs, examples, or behavior improvements.
mod-backend
cuDNN backend API, graph execution, descriptors, engines, or backend integration.
orig-nv-eng
Reported or requested by NVIDIA engineering.
Add MoE + expert-parallel Python API surface, PyTorch reference, and tests
#389
opened Jul 14, 2026 by
Anerudhan
Collaborator
Loading…
DSA: Add FP8/MXFP8 and compressed Top-K indexer paths
#370
opened Jul 9, 2026 by
jiayus-nvidia
Contributor
•
Draft
[DRAFT] Add framework-agnostic operator APIs with JAX support
#363
opened Jul 8, 2026 by
mgoldfarb-nvidia
•
Draft
benchmark: compute causal SDPA FLOP counts without allocating masks
#347
opened Jul 5, 2026 by
fallintoplace
Contributor
Loading…
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.