Skip to content

Pull requests: NVIDIA/TransformerEngine

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[PyTorch] Add head-parallel FA4 backward for cuDNN CP attention attention community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3510 opened Sep 11, 2026 by bzantium Loading…
8 of 13 tasks
[PyTorch] Fuse MoE chunk sorting and padding community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3509 opened Sep 11, 2026 by bzantium Loading…
7 of 13 tasks
[Draft] Port cuDNN frontend attention to Python API
#3508 opened Sep 11, 2026 by vcherepanov-nv Collaborator Draft
13 tasks
[PyTorch] Avoid temporary state casts when loading FusedAdam checkpoints community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3507 opened Sep 11, 2026 by bzantium Draft
5 of 13 tasks
Fix THD P2P pad detection and tail-zero gating community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3506 opened Sep 10, 2026 by RPalmr Loading…
[PyTorch] Support FP8 weight caching in compiled Linear
#3505 opened Sep 10, 2026 by pggPL Collaborator Loading…
6 of 8 tasks
[Pytorch][Attention] Fix FP8 THD backward workspace scaling
#3504 opened Sep 10, 2026 by sudhakarsingh27 Member Loading…
4 of 13 tasks
MOE Sequential Block with Dispatch and Combine as Basic Ops 2.20
#3503 opened Sep 10, 2026 by vthumbe1503 Collaborator Loading…
13 tasks
[PyTorch] Add timestep-conditioned AdaptiveLayerNorm to op fuser community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3501 opened Sep 10, 2026 by zupengwang Loading…
8 of 9 tasks
Fix hybrid extra state size mismatch
#3499 opened Sep 9, 2026 by negvet Collaborator Draft
13 tasks
[JAX] Fix undefined sr_rng_state that breaks --dry-run in two encoder examples community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3498 opened Sep 9, 2026 by Anai-Guo Contributor Loading…
[PyTorch] torch.compile for an OperationFuser group holding one operation
#3496 opened Sep 8, 2026 by pggPL Collaborator Loading…
7 of 13 tasks
[Docs] Add Mixture of Experts guide documentation Improvements or additions to documentation
#3494 opened Sep 7, 2026 by pggPL Collaborator Loading…
[PyTorch] [torch.compile] Prepare LayerNormLinear and LayerNormMLP for torch.compile
#3493 opened Sep 7, 2026 by pggPL Collaborator Loading…
6 of 8 tasks
[PyT] Disable FA3 for training when head_dim_qk != head_dim_v community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3490 opened Sep 7, 2026 by yuweih205 Loading…
6 tasks done
Add native transport for dynamic context parallelism community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3489 opened Sep 6, 2026 by xiaoyao0115 Draft
Support row-only MXFP8 distributed master-weight casts community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3488 opened Sep 6, 2026 by xiuhu17 Contributor Loading…
5 tasks done
[PyTorch] Split FusedAttnFunc into single-argument forward/backward helpers 2.20
#3480 opened Sep 4, 2026 by pggPL Collaborator Loading…
6 of 13 tasks
Drop the unsupported window_size argument from attention backend queries community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3479 opened Sep 4, 2026 by Anai-Guo Contributor Loading…
[PyT] Linear Attention API 2.20
#3477 opened Sep 4, 2026 by KshitijLakhani Collaborator Draft
13 tasks
[PyTorch] Enable fused activation recompute for ScaledTanhSReLU community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3473 opened Sep 3, 2026 by wanyingw Contributor Draft
13 tasks
[PyTorch] torch.compile support for FusedAttention
#3472 opened Sep 3, 2026 by pggPL Collaborator Draft
7 of 13 tasks
[PyTorch] DeepSeekV3Layer: full MoE transformer layer (MLA + DeepSeek MoE)
#3471 opened Sep 3, 2026 by pggPL Collaborator Loading…
7 of 13 tasks
[All] Guard THD learnable dSink on older cuDNN attention bug Something isn't working
#3470 opened Sep 3, 2026 by KshitijLakhani Collaborator Loading…
2 of 13 tasks
[PyTorch] Fix FP8 illegal memory access in single-process multi-GPU execution community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3469 opened Sep 3, 2026 by SuperGoodGame Loading…
ProTip! Exclude everything labeled bug with -label:bug.