跳到正文
原文
X:Unsloth (@UnslothAI)· @analogalok·· 2026-08-27AI 评分62

单张 24GB RTX 4090 跑 Qwen 3.8 Flash Next 125B MoE,25 万上下文达 21 tokens/s

RT by @UnslothAI: The VRAM barrier is officially dead. I just ran Qwen 3.8 Flash Next (MoE) 125B A6B with a 250,000 context window on a single 24GB RTX 4090. 21 tokens/sec decode. 364 t/s prefill. no mtp. no dflash. no kv cache quantization! We are running datacenter models on consumer hardware. Tested on Ubuntu 22 | CUDA 13.0 | PCIe 4.0 x16 | 110 GB DDR4 System RAM with a continuous 28k prompt across all runs. ### The Benchmarks & Scaling # 1. Hybrid Offload (-ncmoe 40 @ 80k Context) Offloaded 40 expert layers to the GPU, pushing VRAM to the ceiling. ./build/bin/llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 80000 --port 8080 -v --fit off -b 4096 -ub 4096 -ncmoe 40 Prefill: 383.85 t/s | Decode: 22.52 t/s Footprint: 23.85 GB VRAM | 97 GB RAM # 2. Full CPU MoE Offload (-cmoe @ 80k Context) Pinned all 512 expert layers to DDR4 RAM (-cmoe), keeping attention on the 4090. llama.cpp flags: (Same as above, replace -ncmoe 40 with -cmoe) Prefill: 355.72 t/s | Decode: 20.8

AI 导读

作者在单张 24GB RTX 4090 上运行 Qwen 3.8 Flash Next(MoE)125B A6B,25 万上下文窗口下解码 21 tokens/s、prefill 364 t/s,未使用 mtp、dflash 和 kv cache 量化。

来源:X:Unsloth (@UnslothAI) · x.lingyaoai.com