跳到正文
原文
0xBakeer· @0xBakeer · X·· 10 天前AI 评分45

单台DGX Spark多模型推理速度实测

AI 导读

这周大家都在晒 3、4 台 Spark 集群。这是一台 DGX Spark 为单个用户跑出的成绩,全部在我的机器上实测: 开启 OpenBMB 的 drafter MiniCPM5-2B:100.8 tok/s Qwen3.6-35B-A3B:89.7 Ling-3.0-flash:69.5 Qwen3.8-Flash-Next:55.5 DeepSeek-V4-Flash:47.9(EXL3) Gemma-4-E2B:39.2 Qwen3.8-27B:33.4(EXL3 + DFlash2) Tinfield-1 2-bit:30.7 一次跑一个模型,256 tokens 输入,256 输出

正文 · 原文

Everyone is posting 3 and 4 Spark clusters this week. Here is what ONE DGX Spark does for one user, all measured on my box:

MiniCPM5-2B, OpenBMB's drafter on: 100.8 tok/s

Qwen3.6-35B-A3B: 89.7

Ling-3.0-flash: 69.5

Qwen3.8-Flash-Next: 55.5

DeepSeek-V4-Flash: 47.9 in EXL3

Gemma-4-E2B: 39.2

Qwen3.8-27B: 33.4 EXL3 + DFlash2

Tinfield-1 at 2-bit: 30.7

One model at a time, 256 tokens in, 256 out

来源:0xBakeer · x.com