提出主动反思式上下文管理框架,解决长任务中信息漂移问题。
ARC: Active and Reflection-driven Context Management for Long-Horizon Information Seeking Agents
- 通过反思监测与动态修正,让上下文随推理过程持续优化。
- 在中文长任务数据集上提升11%准确率,优于传统压缩方法。
- 适合需要长期推理的智能搜索、科研助手等场景。
大语言模型在深度搜索和长周期信息获取任务中表现日益重要,但随着交互历史增长,性能常因上下文漂移而下降,表现为内部状态失准。现有方法多依赖原始积累或被动摘要,将上下文视为静态对象,导致早期错误持续累积。为此,我们提出ARC框架,首次将上下文管理建模为一个主动、反思驱动的过程,将其视作执行过程中的动态推理状态。ARC通过反思驱动的监控与修订机制,在检测到偏差或退化时主动重构工作上下文。在多个挑战性长周期信息获取基准测试中,ARC稳定超越被动压缩方法,在BrowseComp-ZH数据集上使用Qwen2.5-32B-Instruct时,准确率最高提升11%。
原文摘要 · Abstract (English)
Large language models are increasingly deployed as research agents for deep search and long-horizon information seeking, yet their performance often degrades as interaction histories grow. This degradation, known as context rot, reflects a failure to maintain coherent and task-relevant internal states over extended reasoning horizons. Existing approaches primarily manage context through raw accumulation or passive summarization, treating it as a static artifact and allowing early errors or misplaced emphasis to persist. Motivated by this perspective, we propose ARC, which is the first framework to systematically formulate context management as an active, reflection-driven process that treats context as a dynamic internal reasoning state during execution. ARC operationalizes this view through reflection-driven monitoring and revision, allowing agents to actively reorganize their working context when misalignment or degradation is detected. Experiments on challenging long-horizon information-seeking benchmarks show that ARC consistently outperforms passive context compression methods, achieving up to an 11% absolute improvement in accuracy on BrowseComp-ZH with Qwen2.5-32B-Instruct.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。