跳到正文
原文
X:Unsloth (@UnslothAI)· @UnslothAI·· 2026-09-02AI 评分62

Unsloth 为 Qwen3.8-Flash 加入 MTP,本地推理最高提速 1.7 倍

Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change. GGUFs: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF Guide: https://unsloth.ai/docs/models/qwen3.8-next

AI 导读

Unsloth 宣布 Qwen3.8-Flash 借助 MTP 在本地运行最高提速 1.7 倍,GGUF 在 RTX PRO 6000 上可达 170 tokens/s。

来源:X:Unsloth (@UnslothAI) · x.lingyaoai.com