arXiv:2604.07791cs.AIcs.LG2026-04ACL被引 5

让智能体通过结构化记忆联合优化策略与工具,实现高效自进化学习。

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents

论文配图:SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
图 1 · 摘自论文原文
  • 构建规划与执行融合的结构化经验记忆
  • 利用轨迹关联增强稀疏奖励信号,提升学习效率
  • 适合资源受限环境下需持续进化的智能体应用

近年来基于可验证奖励的强化学习(RLVR)在单轮推理任务中展现出显著潜力。随着自进化代理学习范式的兴起,模型被期望通过合成工具或积累显式经验来学习。然而,现有方法通常依赖大规模大语言模型或多智能体框架,难以部署于资源受限环境。结果导向奖励的固有稀疏性也带来挑战,因为智能体通常仅在任务完成后获得反馈。为此,我们提出基于工具-记忆的自进化代理框架SEARL。不同于直接使用交互经验的方法,本方法构建融合规划与执行的结构化经验记忆,提供一种新的状态抽象,促进在相似场景下的泛化,如工具复用。由此,智能体从历史数据中提取显式知识,并利用跨轨迹相关性密度化奖励信号。我们在知识推理和数学任务上评估该框架,证明其在实现更实用、高效的自我学习方面具有有效性。

原文摘要 · Abstract (English)

Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have demonstrated significant potential in single-turn reasoning tasks. With the paradigm shift toward self-evolving agentic learning, models are increasingly expected to learn from trajectories by synthesizing tools or accumulating explicit experiences. However, prevailing methods typically rely on large-scale LLMs or multi-agent frameworks, which hinder their deployment in resource-constrained environments. The inherent sparsity of outcome-based rewards also poses a substantial challenge, as agents typically receive feedback only upon completion of tasks. To address these limitations, we introduce a Tool-Memory based self-evolving agentic framework SEARL. Unlike approaches that directly utilize interaction experiences, our method constructs a structured experience memory that integrates planning with execution. This provides a novel state abstraction that facilitates generalization across analogous contexts, such as tool reuse. Consequently, agents extract explicit knowledge from historical data while leveraging inter-trajectory correlations to densify reward signals. We evaluate our framework on knowledge reasoning and mathematics tasks, demonstrating its effectiveness in achieving more practical and efficient learning.

自进化强化学习智能体经验记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。