[perf] MiniMax-H3 LoRA training on a B200 node: 6.32 s to 2.53 s at 8 GPUs, mostly from the attention kernels - #1680
Draft
TarzanZhao wants to merge 3 commits into
Draft
TarzanZhao wants to merge 3 commits into
TarzanZhao wants to merge 3 commits into