跳到正文
原文
X:Rohan Paul (@rohanpaul_ai)· @rohanpaul_ai·· 20 小时前AI 评分64

NVIDIA 论文:长任务中模型准确率下降,建议给条目编号并分批处理

RT by @rohanpaul_ai: New Nvidia paper: AI models get sloppier as jobs get longer, even inside their context window, so number every item and split big jobs into small chunks. Model size didn't guarantee reliability on long, repetitive jobs Picture an agent updating a huge invoice file line by line. It can read the whole file and still skip a line or update the wrong record. NVIDIA tested 7 open models on simple, repetitive jobs like adding numbers and sorting lists. Average accuracy was 62.8% lower on 128K-token jobs than on 4K-token jobs. Even the best model got every item right in only 17.1% of the longest jobs. The models seemed to understand the task but lost their place, especially when items had no ID numbers. If your agent works through long lists, give every item an ID, process them in small batches, and check every line of output. – arxiv. org/abs/2609.38712 Title: "Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability"

AI 导读

NVIDIA 新论文测试 7 个开源模型在加数字、排序列表等简单重复任务上的表现,128K token 任务的平均准确率比 4K token 任务低 62.8%,最好的模型在最长的任务中也只有 17.1% 全部答对。

来源:X:Rohan Paul (@rohanpaul_ai) · x.lingyaoai.com