提出自进化智能体框架,让机器学会从经验中持续学习与成长。
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

- 采用快慢双循环机制,通过在线蒸馏实现经验的读取、使用、书写与管理。
- 在多领域测试中性能超越现有记忆系统11.5%,接近3970亿参数大模型。
- 适合研究智能体自我进化、长期记忆管理及高效训练方法的学者参考。
记忆已成为自进化智能体的标准基础,但保存经验不等于掌握进化能力。现有记忆系统可存储轨迹、检索反思或积累技能,却常缺乏选择有用经验、应用知识、生成可复用认知并维护增长知识库的综合能力。本文提出OPD-Evolver,一种基于在线策略自蒸馏的慢-快协同进化框架。快速循环中,智能体通过四级记忆层级实现经验的读取、利用、写入与维护,支持测试时快速进化;慢速循环中,基于结果校准的记忆归因与特权回顾蒸馏出四项核心能力至可部署策略。在多领域基准测试中,OPD-Evolver性能超越ReasoningBank最高达11.5%,优于训练型方法Skill0约5.8%。分析表明,该模型内化了高价值经验与记忆管理机制,使OPD-Evolver-9B可挑战如Qwen3.5-397B-A17B和Step-3.5-Flash等巨型模型,推动记忆增强型智能体迈向真正具备进化能力的智能体演进者。
原文摘要 · Abstract (English)
Memory has become a standard substrate for self-evolving agents, yet retaining experience is not the same as learning how to evolve through it. Existing memory agents can store trajectories, retrieve reflections, or accumulate skills, but often lack the holistic competence to select useful experience, act on it, write reusable knowledge, and maintain a growing repository. We introduce OPD-Evolver, a slow-fast co-evolution framework that cultivates such an agent evolver through on-policy self-distillation. In the fast loop, OPD-Evolver interacts with a four-level memory hierarchy to read, use, write, and maintain experience for rapid test-time evolution. In the slow loop, outcome-calibrated memory attribution and privileged hindsight distill these four abilities into the deployable policy. Across multi-domain benchmarks, OPD-Evolver surpasses memory systems such as ReasoningBank by up to 11.5%, and training-based methods such as Skill0 by ~5.8%. Further analysis shows that OPD-Evolver internalizes high-value experience and memory management, enabling OPD-Evolver-9B to challenge giant counterparts such as Qwen3.5-397B-A17B and Step-3.5-Flash, pointing beyond memory-augmented agents toward genuinely qualified agent evolvers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。