arXiv:2505.11942cs.AI2025-05被引 51

首个评估大模型代理持续学习能力的基准,推动智能体长期记忆进化

LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners

  • 构建三类交互环境下的技能关联任务,支持自动验证与模块扩展
  • 实测发现传统经验回放对大模型代理效果有限,因信息无关且上下文受限
  • 提出群体自一致机制,显著提升代理长期学习表现,适合研究持续学习者

持续学习对动态环境中运行的智能体至关重要。当前基于大语言模型(LLM)的代理仍为无状态系统,无法积累或迁移知识。现有基准将代理视为静态系统,无法评估其持续学习能力。我们提出LifelongAgentBench,首个统一的基准,用于系统性评估LLM代理的持续学习能力。该基准在三个交互环境(Database、Operating System、Knowledge Graph)中提供基于技能、相互依赖的任务,支持自动标签验证、可复现性与模块化扩展。大量实验表明,传统经验回放对LLM代理效果有限,主要由于信息无关和上下文长度限制。我们进一步提出群体自一致机制,显著提升持续学习性能。期望LifelongAgentBench能推动具备适应性与记忆能力的LLM代理发展。

原文摘要 · Abstract (English)

Lifelong learning is essential for intelligent agents operating in dynamic environments. Current large language model (LLM)-based agents, however, remain stateless and unable to accumulate or transfer knowledge over time. Existing benchmarks treat agents as static systems and fail to evaluate lifelong learning capabilities. We present LifelongAgentBench, the first unified benchmark designed to systematically assess the lifelong learning ability of LLM agents. It provides skill-grounded, interdependent tasks across three interactive environments, Database, Operating System, and Knowledge Graph, with automatic label verification, reproducibility, and modular extensibility. Extensive experiments reveal that conventional experience replay has limited effectiveness for LLM agents due to irrelevant information and context length constraints. We further introduce a group self-consistency mechanism that significantly improves lifelong learning performance. We hope LifelongAgentBench will advance the development of adaptive, memory-capable LLM agents.

持续学习大模型代理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。