跳到正文
原文
X:MiniMax (@MiniMax_AI)· @MiniMax_AI·· 15 天前AI 评分40

MiniMax-H3 获 VC-Attention 免训练低比特加速,B200 上提速 1.6 倍

Three paths to faster video attention: compute the same interactions more efficiently, compute fewer in full, or change how information is mixed. Here’s a visual guide. 👇 Thanks to Nunchux AI and collaborators for VC-Attention, bringing training-free low-bit acceleration to MiniMax-H3, with better fidelity than SageAttention2 in the B200 evaluation. The approach balances speed and fidelity: V-Smooth reduces value quantization error, while ExpCast-FP8 makes softmax faster through approximation. Excited to see the community keep building on H3. Could combining low-bit computation with sparse methods like Sol-Attn push efficiency further? We’re looking forward to seeing that explored.

AI 导读

Nunchux AI 及合作者推出 VC-Attention,为 MiniMax-H3 带来免训练低比特注意力加速,在 B200 上较 FlashAttention-4 提速 1.6 倍、B300 上提速 1.5 倍,保真度优于 SageAttention2。

来源:X:MiniMax (@MiniMax_AI) · x.lingyaoai.com