dair_ai· @dair_ai · X·· 4 天前AI 评分
NVIDIA 论文量化长程智能体在长上下文下的准确率衰减
DAIR.AI 介绍 NVIDIA 关于长程智能体的论文:模型即使接受 128K tokens 上下文,任务执行越久越容易出错。
Banger paper from NVIDIA on long running agents.
A model can accept 128K tokens of context and still make more mistakes the longer it works through a task.
If your agent loses its place partway through a long table or ledger, this work measures what causes it.
The setup:
Long-Transduction asks a model to keep reading, updating and outputting state-dependent results over thousands of outputs, and varies three factors separately.
Results:
Across seven open-weight models, accuracy drops 62.8% when context grows from 4K to 128K, 36.5% when only the input format changes, and 39.9% when the per-step operation gets harder.
Paper: academy.dair.ai/papers/stayi…
来源:dair_ai · x.com