vllm_project· @vllm_project · X·· 1 天前AI 评分
vLLM 使 DeepSeek-V4.1-Flash 推理提速 1.9 倍、吞吐提升 5.3 倍
vLLM 官方宣布 DeepSeek-V4.1-Flash 在 vLLM 上运行三周后,低并发下速度提升 1.9 倍,在 SemiAnalysis AgentX 上每用户 150 TPS 时吞吐达到 5.3 倍。官方提供了可逐步查看的交互式图表,详见 vllm.ai/blog/2026-10-07-deep…。
1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS per user on @SemiAnalysis_ AgentX.
Here is how, with interactive figures you can step through 🧵 vllm.ai/blog/2026-10-07-deep…
来源:vllm_project · x.com