X:Unsloth (@UnslothAI)· @UnslothAI·· 2026-09-02AI 评分
Unsloth 为 Qwen3.8-Flash 加入 MTP,本地推理最高提速 1.7 倍
Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change. GGUFs: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF Guide: https://unsloth.ai/docs/models/qwen3.8-next
Unsloth 宣布 Qwen3.8-Flash 借助 MTP 在本地运行最高提速 1.7 倍,GGUF 在 RTX PRO 6000 上可达 170 tokens/s。
来源:X:Unsloth (@UnslothAI) · x.lingyaoai.com