压缩上下文会削弱近期交互影响,导致智能体行为不稳定。
Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability

- 用配对闭环实验验证压缩事件的影响,保持模型不变
- 在AppWorld上提升任务表现与多轮运行可靠性
- 适合关注长时序智能体稳定性的研究者
循环上下文压缩虽能控制长时序智能体的上下文增长,但其行为影响尚不明确。本初步实证研究发现,压缩会减弱近期交互的影响,导致动作阻塞、重复探索及跨运行不稳定性。为此,我们提出TRACE框架,通过从同一环境状态出发的配对闭环延续,评估单个压缩事件,并利用摘要偏好优化自然语言压缩提示,同时保持所有模型冻结。在AppWorld上的初步结果表明,该方法在任务性能、多轮运行可靠性及上下文-执行效率方面均优于现有压缩基线。这些发现为边界局部评估作为可靠智能体上下文压缩的可行方向提供了早期证据。
原文摘要 · Abstract (English)
Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood. In this preliminary empirical study, we show that compression can weaken the influence of recent interactions, increasing blocked actions, repeated exploration, and instability across runs. Motivated by these observations, we introduce TRACE, a verifier-guided framework that evaluates individual compaction events through paired closed-loop continuations from the same environment state and uses summary preferences to optimize a natural-language compression prompt while keeping all models frozen. Initial results on AppWorld show improvements over existing compression baselines in task performance, multi-run reliability, and context--execution efficiency. These findings provide early evidence for boundary-local evaluation as a promising direction for reliable agent context compression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。