跳到正文
原文
rohanpaul_ai· @rohanpaul_ai · X·· 19 小时前AI 评分61

上海AI实验室论文:用验证过的 Agent 技能运行数据训练模型提升 Agent 能力

AI 导读

上海AI实验室新论文提出 SkillGym,把每个技能文件转成带代码检查器的沙箱任务,只用通过校验的运行数据微调模型,使模型在不加载技能文件的情况下也变得更强。

正文 · 原文

New Shanghai AI Laboratory paper finds that training on verified runs of human-written agent skills makes a model a better agent, even without the skill files.

Prompt-time skills depend on retrieval and instruction following, and fine-tuning on verified skill runs reduces that dependence.

Turning each skill file into a sandboxed task with a pass-or-fail checker produced training data that lifted Terminal-Bench 2.1 success by 19.10 points in Claude Code.

Skill files usually sit in the prompt, so they only help if the agent finds and follows them.

SkillGym turns each skill into a sandboxed task with a code checker, then trains on the runs that pass.

In Claude Code, Qwen3.5-35B-A3B jumped from 39.33% to 58.43% on Terminal-Bench 2.1. With no skill files, it scored 26.81% on SkillsBench, beating the base model with skills at 23.34%.

Loading the skills on top still helps, lifting it to 51.47%.

来源:rohanpaul_ai · x.com