提出路径依赖评估框架,揭示长期学习中经验积累对智能体表现的影响。
PATH-Bench: Path-Dependent Evaluation of Lifelong Agents

- 构建路径依赖评估框架,控制历史经验的正负影响
- 发现强迁移不等于高保留,后期经验可重塑早期成果
- 提出选择性经验使用机制,显著减少遗忘并提升迁移
长期语言模型智能体通过外部学习状态存储过往交互记忆或可复用技能,但现有基准很少考虑累积经验路径如何影响知识传递与保留。本文提出PATH-Bench,一个用于路径依赖评估的基准。该基准通过多模型上下文学习估计任务间的有向关系,构建带有受控帮助与干扰历史的探针序列,并重复评估探针任务以测量平均性能、前向迁移、反向迁移和遗忘率。在单轮代码生成和多轮工具使用任务上,评估了八种代表性智能体在正向与负向主导历史下的表现。结果表明,经验效用取决于经验表征方式与任务交互结构的共同作用;强迁移并不保证保留;后期经验可重塑早期学习成果。基于此,提出选择性经验使用(SEU)机制,调控路径积累经验对新任务的影响,筛选有益项并过滤干扰。SEU在多数场景下持续降低遗忘率并提升前向迁移能力。PATH-Bench不仅提供受控评估框架,也为设计更精准、鲁棒的长期智能体提供实践指导。
原文摘要 · Abstract (English)
Lifelong LLM agents increasingly adapt through external learning states that store past interactions as retrievable memories or reusable skills, yet existing benchmarks rarely account for how the path of accumulated experience shapes what agents transfer and retain. In this work, we establish PATH-Bench, a benchmark for path-dependent evaluation of lifelong agents. PATH-Bench estimates directed task relationships via multi-model in-context learning, constructs probe-centered sequences with controlled helpful and interfering histories, and repeatedly evaluates probe tasks to measure average performance, forward transfer, backward transfer, and forgetting. We evaluate eight representative agents on single-turn code generation and multi-turn tool-use tasks under positive- and negative-dominant histories. Benchmark results show that experience utility depends jointly on how experience is represented and on the task's interaction structure, that strong transfer does not ensure retention, and that later experience can reshape gains acquired earlier in the learning path. Based on these findings, we propose Selective Experience Use (SEU), an agent harness that regulates how path-accumulated experience influences each new task, admitting helpful items while filtering out potential interference. SEU consistently reduces forgetting while improving forward transfer in the majority of settings. The PATH-Bench provides both a controlled evaluation framework and actionable guidance for designing more selective and robust lifelong agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。