arXiv:2609.04128cs.AI2026-09

通过离线迭代升级环境难度,持续提升智能体训练效果。

Environment Evolution for Terminal Agents

  • 离线演化环境,按生成代数逐步增加难度
  • 在终端基准测试中使模型性能提升14.4~18.0个百分点
  • 适合需要长期强化学习的复杂任务训练场景

扩展可交互且可验证的环境对训练终端智能体至关重要。随着前沿模型能力提升,从头合成的环境逐渐变得不具挑战性,学习信号有限。现有共演化方法依赖在线推演,在模型变强后难以持续提供有效学习信号。本文提出环境演化机制,通过离线方式增量提升环境难度,并在训练过程中按代数调度演化环境,实现持续学习信号供给。基于多轮学习目标,我们推导出三种影响环境难度的方向,并通过环形工程的多智能体框架实现演化。在Hy4 preview、Claude Opus 5和GPT-5.6 Sol上的定量推演实验表明,环境演化能持续生成更难环境。在Qwen3.6-27B和Qwen3.6-35B-A3B上进行简单长时程强化学习验证,其在Terminal-Bench 2.1上的性能分别提升14.4和18.0个百分点。

原文摘要 · Abstract (English)

Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evolution methods iteratively synthesize environments near the model's learnable frontier based on weaknesses exposed during rollouts. However, their dependence on on-policy rollouts limits generalization and the continuous provision of learning signals as the model becomes stronger. In this paper, we propose environment evolution, which incrementally increases environment difficulty off-policy and schedules the evolved environments generation by generation during training to provide continuous learning signals. We derive three evolution directions that influence environment difficulty from the multi-turn learning objective and then implement evolution along these directions through a loop-engineered multi-agent harness. Quantitative rollout experiments with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show that environment evolution consistently produces more difficult environments. We validate its effectiveness on Qwen3.6-27B and Qwen3.6-35B-A3B through simple long-horizon RL training, improving their performance by 14.4 and 18.0 percentage points on Terminal-Bench 2.1, respectively.

环境演化强化学习智能体训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。