-
Notifications
You must be signed in to change notification settings - Fork 2.6k
Pull requests: NVIDIA/TensorRT-LLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[None][perf] Add BSX multi-tier CuTe DSL top-k decode kernels (stacked on #16457)
#16877
opened Jul 26, 2026 by
longcheng-nv
Collaborator
•
Draft
5 of 7 tasks
[None][perf] optimize native V2 KV event production
#16876
opened Jul 26, 2026 by
alec-flowers
Collaborator
•
Draft
[None][perf] Address inter-iter idle times
#16875
opened Jul 26, 2026 by
brb-nv
Collaborator
Loading…
1 task done
[None][perf] serve: always use msgspec msgpack for disagg orchestrator->worker body
#16873
opened Jul 25, 2026 by
Tabrizian
Member
Loading…
1 task done
[https://nvbugs/6480621][feat] add disaggregated lifecycle diagnostics
#16872
opened Jul 25, 2026 by
chienchunhung
Collaborator
•
Draft
[https://nvbugs/6506918][fix] Reject the incompatible configuration up front in
Linear.__init__ with an…
#16870
opened Jul 25, 2026 by
trtllm-agent
Collaborator
Loading…
2 tasks done
[None][feat] publish native V2 KV cache events
#16869
opened Jul 25, 2026 by
alec-flowers
Collaborator
•
Draft
[https://nvbugs/6507114][fix] Extend the per-model qwen3.5_moe_35b.yaml with…
#16868
opened Jul 25, 2026 by
trtllm-agent
Collaborator
Loading…
2 tasks done
[https://nvbugs/6240584][fix] Qwen3ToolParser: bare-JSON fallback for reasoning-preceded tool calls
#16866
opened Jul 25, 2026 by
JunyiXu-nv
Collaborator
Loading…
1 task done
[None][perf] Fuse MiniMax-M3 MoE routing
#16859
opened Jul 25, 2026 by
peihu-nv
Collaborator
Loading…
1 task done
[https://nvbugs/6506920][fix] Add
stderr, stderr_margin_sigmas fields to HypothesisTestingParams…
#16858
opened Jul 25, 2026 by
trtllm-agent
Collaborator
Loading…
2 tasks done
[None][perf] Avoid Index-K cache materialization for MSA
#16856
opened Jul 25, 2026 by
peihu-nv
Collaborator
Loading…
1 task done
[None][test] Stabilize scaffolding OpenAI worker tests
#16853
opened Jul 24, 2026 by
Mgluhovskoi
Contributor
•
Draft
[None][perf] Fuse MiniMax-M3 prefill projections
api-compatible
Accepted LLM API contract change that is backwards-compatible
[None][fix] Increase max top logprobs limit
#16851
opened Jul 24, 2026 by
yibinl-nvidia
Collaborator
Loading…
1 task
[https://nvbugs/6506990][fix] Don't treat NVLE-only as CC enabled
#16850
opened Jul 24, 2026 by
dhansen-nvidia
Collaborator
Loading…
1 task done
[None][perf] fp8 block scale quant fusion in SM90 Cutlass MoE
#16849
opened Jul 24, 2026 by
amukkara
Collaborator
Loading…
1 task done
[None][fix] Update VisualGen test CODEOWNERS
#16848
opened Jul 24, 2026 by
yibinl-nvidia
Collaborator
Loading…
1 task done
Previous Next
ProTip!
Filter pull requests by the default branch with base:main.