Chollet:RLVR 或只推高数学与代码能力
Francois Chollet 提出假设:AI 的"锯齿状前沿"可能主要集中在数学与代码,可通过 RLVR 无限推进,而其他领域因受限于人类生成数据将趋于平台期。他指出非可验证领域的模型性能虽仍在稳步提升,但远慢于数学与代码,其提升究竟源于 RLVR 带来的更高 G,还是仅取决于持续注入的新人类数据,尚无定论。这一问题的答案将影响 AI 发展的诸多判断。
What if the jagged frontier is mainly math + code (which you can push arbitrarily far with RLVR), and everything else starts to plateau because it is still bottlenecked by human generated data?
Model performance in non-verifiable areas has kept improving steadily, albeit much slower than for math and code. But is that steady improvement a side effect of a higher G (itself driven by RLVR), or only a function of the amount of new human data getting injected into training (which is still continually happening on a massive scale)?
A lot of things depend on the answer to this question
来源:fchollet · x.com