跳到正文
原文
X:面壁智能 OpenBMB (@OpenBMB)· @OpenBMB·· 4 天前AI 评分41

清华NLP等提出 One-Shot OPD:训练集缩到一条 query,仍拿下全量 OPD 87% 的增益

Post-training pipelines now use on-policy distillation (OPD) to hand a student the teacher's full next-token distribution at every prefix it visits: the student samples its own rollouts, and Qwen3, MiMo, GLM-5, DeepSeek-V4 and Kimi K3 all pair it with SFT and RL. Yet work on OPD has almost all stood on the algorithm side, treating the training set as given—so how much of OPD's gain does the data account for? Introducing One-Shot OPD, from @TsinghuaNLP (OpenBMB member) with the University of Chinese Academy of Sciences, Northeastern University, UIUC and Johns Hopkins University. It cuts the training set to one query, and the answer is that OPD is data-overfed but algorithm-starved. 1️⃣ One query, hundreds of steps. On math it goes from 59.1 to 68.5 by step 300, against 69.8 for full-data OPD—87% of its gain. It holds across code, instruction following and agentic tool use, and across Qwen, Llama and OLMo; a query the student never solves works about as well as one it always solves. 2️⃣

AI 导读

清华NLP(OpenBMB 成员)联合中科院大学、东北大学、UIUC 和约翰霍普金斯大学提出 One-Shot OPD,把 on-policy distillation 的训练集压缩到一条 query。

来源:X:面壁智能 OpenBMB (@OpenBMB) · x.lingyaoai.com