SGLang 适配 NVIDIA Vera Rubin 并加速 Kimi K3 推理
SGLang 官方宣布已将 SGLang 带到 NVIDIA Vera Rubin 早期访问硬件上,与 NVIDIA 合作优化 attention、MoE 和投机验证内核以加速 Kimi K3 推理。
We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference.
Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware.
Highlights:
• Up to 20% faster FP8 MLA at batch 1 / 128K context
• 20% faster KDA verification, with bitwise-identical output
• 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step
SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.
Full results and engineering details 👉 lmsys.org/blog/2026-10-09-ve…
来源:sgl_project · x.com