Meta Superintelligence Labs 论文提出 Sharpening Tax,指出 RL 后训练削弱基座模型的测试时可扩展性
Meta Superintelligence Labs 的论文发现,在采样量充足时,配轻量框架的基座模型往往比 RL 后训练版本解决更多智能体任务:后训练模型在 pass@1 上占优,但在 BFCL v4 multi-turn、ACEBench 和 WebShop 上,大 K 时基座模型常能解决后训练模型始终解决不了的任务。
当前显示原文,中文译文尚未提供。
Banger paper from Meta Superintelligence Labs.
They find something super interesting and unexpected.
(bookmark it)
Base models with a light harness often solve more agentic tasks than their RL post-trained versions when both get enough samples.
Post-trained models win on pass@1. At large K, base models frequently solve tasks the post-trained ones never solve on BFCL v4 multi-turn, ACEBench, and WebShop.
This is because post-training pushes each task toward always solved or never solved. Consistency goes up, and coverage goes down.
The authors call the lost test-time scalability the Sharpening Tax. Across 42 base and post-trained pairs, it shows up in most settings, grows with model size, and can be estimated from a few rollouts.
Their fix, PTGS, sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax and also raises pass@1.
Paper: academy.dair.ai/papers/sharp…
来源:dair_ai · x.com