让大模型和训练机制一起进化,提升自主强化学习效果
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning

- 大模型与训练框架协同进化,通过反馈不断优化
- 在长序列软件工程任务中表现超越人工设计基准
- 适合研究自主智能体与强化学习的学者和开发者
自主大模型训练常被视为配方搜索,导致训练框架保持静态。这一局限在智能体强化学习中尤为突出,因瓶颈变化与标量奖励掩盖了多种失败模式。我们提出EvoTrainer,一种自主训练框架,通过实证反馈协同演化大模型策略与训练侧的驾驭机制:它诊断轨迹级证据、修订诊断逻辑、回测干预措施,并积累可复用技能。在数学推理、竞赛编程代码生成及仓库级软件工程任务上评估,EvoTrainer在相同数据、代码库与评估协议下达到或超过人工设计的强化学习基准,尤其在长时序智能体软件工程任务中提升显著。轨迹分析显示,保留策略在不同领域分化,演化诊断防止无效高分分支被选中,可复用技能影响后续搜索方向。自主大模型强化学习应从配方搜索转向策略与训练驾驭机制的联合演化。
原文摘要 · Abstract (English)
Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in agentic RL, where shifting bottlenecks and scalar rewards mask diverse failure modes. We introduce EvoTrainer, an autonomous training framework that co-evolves LLM policies and training-side harnesses through empirical feedback: it diagnoses rollout-level evidence, revises diagnostics, backtests interventions, and accumulates reusable skills. Evaluated on mathematical reasoning, competitive-programming code generation, and repository-level software engineering, EvoTrainer matches or exceeds the human-engineered RL references under the same data, codebase, and evaluation protocol, with the largest gain on long-horizon agentic SWE. Trajectory analyses show that retained strategies diverge across domains, evolving diagnostics prevent invalid high-scoring branches from being promoted, and reusable skills shape later search. Autonomous LLM RL should move beyond recipe search toward joint evolution of policies and the training harnesses that interpret them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。