让智能体和训练环境共同进化,提升工具使用能力。
SEAL: Synergistic Co-Evolution of Agents and Learning Environments

- 智能体与环境闭环协同进化,共享失败诊断信号。
- 400样本下平均提分8.25至26.25,支持跨分布迁移。
- 适合研究自进化智能体、交互式学习的学者参考。
大型语言模型智能体通过交互不断改进,但现有自进化方法通常孤立地优化智能体策略或学习环境。我们识别出这一结构性缺陷为“智能体-环境错配”:智能体能力边界在训练中变化,而提供监督的环境保持静态或耦合较弱。为此提出SEAL框架,实现交互式工具使用智能体的闭环协同进化。SEAL在可执行验证下收集策略轨迹,将失败回放诊断为逐轮失败标签,并以此作为环境与策略双侧演化的共享信号。环境通过暴露更清晰的工具可用性提示、约束信息和恢复反馈来优化训练期学习接口;策略则通过诊断引导的优势重加权进行更新。在分布内与分布外的多轮工具使用评估中,SEAL显著提升低资源智能体学习效果:仅用400个训练样本,即在三个骨干模型上获得8.25至26.25的平均分提升,并表现出正向分布外迁移能力。结果表明,联合优化学习者与其训练期学习基底对构建鲁棒自进化大模型智能体具有重要价值。
原文摘要 · Abstract (English)
Large Language Model (LLM) agents are increasingly improved through interaction, yet most self-evolution methods adapt either the policy or the learning environment in isolation. We identify this structural gap as \emph{Agent-Environment Misalignment}: the agent's capability frontier changes during training, while the environment that provides supervision remains static or only weakly coupled to the agent's revealed failures. We propose SEAL, a closed-loop co-evolution framework for interactive tool-use agents. SEAL collects on-policy trajectories under executable verification, diagnoses failed rollouts into turn-level failure labels, and uses these diagnoses as a shared signal for both environment-side adaptation and model-side policy optimization. The environment evolves its training-time learning interface by exposing clearer tool affordance cues, constraint information, and recovery-oriented feedback, while the policy is updated with diagnosis-guided advantage reweighting. Extensive experiments across in-distribution and out-of-distribution multi-turn tool-use evaluations show that SEAL improves low-resource agent learning: with only 400 training samples, it yields +8.25 to +26.25 average-point gains across three backbones and exhibits positive out-of-distribution transfer. These results demonstrate the value of jointly adapting the learner and its training-time learning substrate for robust self-improving LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。