跳到正文
原文
omarsar0· @omarsar0 · X·· 7 小时前AI 评分58

Elvis Saravia 看好后训练降本趋势:Tinker 长上下文 token 降价利好 agentic RL

AI 导读

Elvis Saravia 表示看好让后训练更易获取的趋势,指出在 agentic RL 中长上下文任务昂贵且难以扩展,而 rollout 占用大部分 token,每一轮都要重读不断增长的上下文。

正文 · 原文

Bullish on this trend of making post-training more accessible.

A new post-training era is upon us.

If you work on agentic RL, long-context tasks (a big focus today) are expensive, inefficient, and don't scale well.

I've been diving into RL envs and evals for long-context tasks, and I can see this being useful.

In agent RL, rollouts use most of the tokens. Every turn re-reads the whole growing context, including tool outputs, files, and earlier turns.

Tinker just cut the price of those tokens. Long-context prefill and sampling now cost the same as short context.

This means that evaluating your trained models on long inputs also gets cheaper. Huge win here.

I believe RL will keep unlocking specialized models that slash the cost of critical agent operations. Cheaper long rollouts make them more practical to build.

Own your intelligence stack!

来源:omarsar0 · x.com