UT Austin 研究:LLM Agent 上下文压缩策略可能拖慢运行速度
DAIR.AI 转发 UT Austin 关于 LLM Agent 上下文压缩的研究,该研究在 SWE-bench Verified 和 Terminal-Bench 上跑了近 35,000 次 agent 运行,分别改变压缩方式、触发时机和删除比例三个决策。
Great overview of context compression in LLM Agents
And one interesting, unexpected finding.
If your compaction policy is tuned to cut tokens, it may be making your agent run slower.
Great study from UT Austin on context compression in coding agents.
They ran nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench and varied three decisions separately. These are how context is compressed, when compression triggers, and how much is removed.
On Terminal-Bench with Qwen, policies that use about a third of the tokens can take 20% to 80% longer than keeping full context.
Step-triggered policies cut the most tokens per step but need 10% to 27% more model calls. Threshold-triggered policies cut tokens by 22% to 55% with call counts close to full context.
Results also differ by model. A policy that works well for Qwen drops Devstral to 38.7% and makes it slower, so measure latency and cost per model before choosing one.
Paper: arxiv.org/abs/2609.32961
Chat with Paper: academy.dair.ai/papers/beyon…
来源:omarsar0 · x.com