跳到正文
原文
dair_ai· @dair_ai · X·· 11 小时前AI 评分61

PAIR 论文分析上下文压缩对长程智能体的损害并改进压缩提示词

AI 导读

DAIR.AI 介绍一篇提出 PAIR 的论文,该方法从同一状态重放智能体、分别在有无压缩下对比,以定位上下文压缩对长程智能体的损害。

正文 · 原文

Context compression is a huge bottleneck for long-running agents.

This work finds that context compression hurts long-horizon agents at a few specific points.

They propose PAIR, which replays the agent from the same state with and without a given compression, instead of comparing whole runs that differ in many random ways.

A typical compression adds a few extra steps. The large drops in success come from a small number of compression events.

The harmful compressions drop task conditions the agent hasn't resolved yet. In one Venmo task, the summary dropped the "only from coworkers" filter and reported the total of all 36 payments as the answer.

Other compressions reduce API specs the agent already read to vague prose, so the agent reopens the docs and logs in again, which adds about five steps.

PAIR then diagnoses what information those compressions dropped and rewrites the matching sections of the compression prompt. The agent, compressor model, and tools stay fixed.

On AppWorld, OfficeBench, and tau-Bench Retail, it gives the most consistent task completion of any compressed method and comes close to running with no compression at all.

Compression also lowers run-to-run reliability before it makes tasks unsolvable, so check consistency across repeated runs in your own evals.

Paper: arxiv.org/abs/2609.36526

Chat with Paper: academy.dair.ai/papers/adapt…

来源:dair_ai · x.com