-
Notifications
You must be signed in to change notification settings - Fork 824
Pull requests: NVIDIA/TransformerEngine
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[PyTorch] Add head-parallel FA4 backward for cuDNN CP attention
attention
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3510
opened Sep 11, 2026 by
bzantium
Loading…
8 of 13 tasks
[PyTorch] Fuse MoE chunk sorting and padding
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3509
opened Sep 11, 2026 by
bzantium
Loading…
7 of 13 tasks
[Draft] Port cuDNN frontend attention to Python API
#3508
opened Sep 11, 2026 by
vcherepanov-nv
Collaborator
•
Draft
13 tasks
[PyTorch] Avoid temporary state casts when loading FusedAdam checkpoints
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
Fix THD P2P pad detection and tail-zero gating
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3506
opened Sep 10, 2026 by
RPalmr
Loading…
[PyTorch] Support FP8 weight caching in compiled Linear
#3505
opened Sep 10, 2026 by
pggPL
Collaborator
Loading…
6 of 8 tasks
[Pytorch][Attention] Fix FP8 THD backward workspace scaling
#3504
opened Sep 10, 2026 by
sudhakarsingh27
Member
Loading…
4 of 13 tasks
MOE Sequential Block with Dispatch and Combine as Basic Ops
2.20
#3503
opened Sep 10, 2026 by
vthumbe1503
Collaborator
Loading…
13 tasks
[PyTorch] Add timestep-conditioned AdaptiveLayerNorm to op fuser
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3501
opened Sep 10, 2026 by
zupengwang
Loading…
8 of 9 tasks
[JAX] Fix undefined sr_rng_state that breaks --dry-run in two encoder examples
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3498
opened Sep 9, 2026 by
Anai-Guo
Contributor
Loading…
[PyTorch] torch.compile for an OperationFuser group holding one operation
#3496
opened Sep 8, 2026 by
pggPL
Collaborator
Loading…
7 of 13 tasks
[Docs] Add Mixture of Experts guide
documentation
Improvements or additions to documentation
#3494
opened Sep 7, 2026 by
pggPL
Collaborator
Loading…
[PyTorch] [torch.compile] Prepare LayerNormLinear and LayerNormMLP for torch.compile
#3493
opened Sep 7, 2026 by
pggPL
Collaborator
Loading…
6 of 8 tasks
[PyT] Disable FA3 for training when head_dim_qk != head_dim_v
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3490
opened Sep 7, 2026 by
yuweih205
Loading…
6 tasks done
Add native transport for dynamic context parallelism
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3489
opened Sep 6, 2026 by
xiaoyao0115
•
Draft
Support row-only MXFP8 distributed master-weight casts
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3488
opened Sep 6, 2026 by
xiuhu17
Contributor
Loading…
5 tasks done
[PyTorch] Split FusedAttnFunc into single-argument forward/backward helpers
2.20
#3480
opened Sep 4, 2026 by
pggPL
Collaborator
Loading…
6 of 13 tasks
Drop the unsupported window_size argument from attention backend queries
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3479
opened Sep 4, 2026 by
Anai-Guo
Contributor
Loading…
[PyT] Linear Attention API
2.20
#3477
opened Sep 4, 2026 by
KshitijLakhani
Collaborator
•
Draft
13 tasks
[PyTorch] Enable fused activation recompute for ScaledTanhSReLU
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
[PyTorch] DeepSeekV3Layer: full MoE transformer layer (MLA + DeepSeek MoE)
#3471
opened Sep 3, 2026 by
pggPL
Collaborator
Loading…
7 of 13 tasks
[All] Guard THD learnable dSink on older cuDNN
attention
bug
Something isn't working
#3470
opened Sep 3, 2026 by
KshitijLakhani
Collaborator
Loading…
2 of 13 tasks
[PyTorch] Fix FP8 illegal memory access in single-process multi-GPU execution
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3469
opened Sep 3, 2026 by
SuperGoodGame
Loading…
Previous Next
ProTip!
Exclude everything labeled
bug with -label:bug.