Elvis Saravia 看好后训练降本趋势:Tinker 长上下文 token 降价利好 agentic RL
Elvis Saravia 表示看好让后训练更易获取的趋势,指出在 agentic RL 中长上下文任务昂贵且难以扩展,而 rollout 占用大部分 token,每一轮都要重读不断增长的上下文。
Bullish on this trend of making post-training more accessible.
A new post-training era is upon us.
If you work on agentic RL, long-context tasks (a big focus today) are expensive, inefficient, and don't scale well.
I've been diving into RL envs and evals for long-context tasks, and I can see this being useful.
In agent RL, rollouts use most of the tokens. Every turn re-reads the whole growing context, including tool outputs, files, and earlier turns.
Tinker just cut the price of those tokens. Long-context prefill and sampling now cost the same as short context.
This means that evaluating your trained models on long inputs also gets cheaper. Huge win here.
I believe RL will keep unlocking specialized models that slash the cost of critical agent operations. Cheaper long rollouts make them more practical to build.
Own your intelligence stack!
来源:omarsar0 · x.com