跳到正文
原文
UnslothAI· @UnslothAI · X·· 2 天前AI 评分59

Unsloth 开源本地训练 Decision 模型的方法,Qwen3.5 0.8B 决策基准准确率提升至 74.3%

AI 导读

Unsloth 宣布用户现在可以在本地训练自己的 Decision 模型,其将 Qwen3.5 0.8B 在 3 个决策基准上的综合准确率从 20.7% 提升到 74.3%,仅需 4GB VRAM。借助开源 Unsloth repo,可把 Qwen3.8、Gemma 4 等任意 LLM 转成决策模型;具体做法是用 Unsloth 和 LoRA(r=64)微调一个 Clef head 训练一个 epoch,将下游准确率从 30–37% 提升到 78%。相关 Guide 和 Notebooks 见 unsloth.ai/docs/basics/train,仓库地址 github.com/unslothai/unsloth。

正文 · AI 翻译

你现在可以在本地训练自己的 Decision 模型,就像 Jev 一样!

我们仅用 4GB 显存,就将 Qwen3.5 0.8B 在 3 个决策基准测试上的综合准确率从 20.7% 提升到了 74.3%。

使用我们开源的 Unsloth 仓库,可以将 Qwen3.8、Gemma 4 等任何 LLM 转换为决策模型。

我们使用 Unsloth 和 LoRA(r=64)配合 Clef 头进行了一个 epoch 的微调,将下游准确率从 30–37% 提升到了 78%。

GitHub:github.com/unslothai/unsloth

指南和笔记本:unsloth.ai/docs/basics/train…

来源:UnslothAI · x.com