跳到正文
原文
sgl_project· @sgl_project · X·· 2 天前AI 评分55

SGLang 适配 NVIDIA Vera Rubin 并加速 Kimi K3 推理

AI 导读

SGLang 官方宣布已将 SGLang 带到 NVIDIA Vera Rubin 早期访问硬件上,与 NVIDIA 合作优化 attention、MoE 和投机验证内核以加速 Kimi K3 推理。

正文 · 原文

We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference.

Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware.

Highlights:

• Up to 20% faster FP8 MLA at batch 1 / 128K context

• 20% faster KDA verification, with bitwise-identical output

• 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step

SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.

Full results and engineering details 👉 lmsys.org/blog/2026-10-09-ve…

来源:sgl_project · x.com