自改进智能体在复杂任务中表现不稳定,顺序和环境描述不清晰是主因。
On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
- 通过多轮实验与任务顺序随机化,发现性能受噪声和顺序影响大
- 引入详细评分标准和环境反馈后,性能提升但仍有明显差距
- 提醒研究者需重视评估严谨性,设计可监督的人机协作系统
基于记忆的自改进智能体——通过在线任务流学习并持续优化自身能力——近年来展现出巨大潜力。然而,这类方法的可靠性问题长期被忽视。本文对两种基于记忆的方法进行了全面重评估,从两个维度扩展测试:(1) 多次运行以量化方差,(2) 随机打乱任务顺序,探究任务顺序的影响。实验揭示两大脆弱性:第一,在复杂环境与多步任务中,评估本身具有内在噪声,而叠加自改进机制会进一步放大噪声;第二,智能体的改进高度依赖任务顺序,先前研究采用的默认顺序隐含了教学课程,构成成功前提。通过人工分析记忆内容,我们推测任务与环境描述不足是根源。通过在记忆构建中加入详细评分标准与环境反馈信息,部分缓解了性能下降,但仍存在显著差距,表明尚有未识别因素导致脆弱性。未来工作应倡导更严格的评估协议,报告多轮结果,并在挑战性条件下进行压力测试。同时,研究呼吁设计支持有效人类监督的系统,避免智能体出现不可预见的失败。
原文摘要 · Abstract (English)
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening the scope of evaluation along two axes: (1) including multiple runs to quantify variance, and (2) randomly shuffling the tasks to investigate the effect of task order. Through these experiments, we make two observations that expose the fragility of current methods: First, agent evaluation is inherently noisy in complex environments and on multi-step tasks, and stacking a self-improving loop on top can further amplify this noise. Second, the agent's improvement is highly dependent on task order. Prior works often adopt default orderings that impose an implicit curriculum, acting as a hidden prerequisite for success. To better understand this fragility, we manually examine the agents' memory and hypothesize that task and environment underspecification contribute to this fragility. We validate this hypothesis by incorporating information that enables better specification, such as detailed rubrics and environment feedback, into the memory construction process. While this added information partially closes the performance degradation in previous experiments, significant gaps still remain, suggesting that other uncharacterized factors contribute to this fragility. Looking ahead, our work advocates for more rigorous evaluation protocols for self-improving agents by reporting results across multiple runs and stress-testing them under challenging conditions. Moreover, our findings on underspecification call for systems and interfaces that enable effective human oversight, preventing agents from failing in unforeseeable ways.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。