arXiv:2512.22857cs.CLcs.AI2025-12被引 10

自动生成高难度可验证任务环境,提升智能体强化学习稳定性与效率

AutoForge: Automated Environment Synthesis for Agentic Reinforcement Learning

  • 构建自动化环境生成流水线,支持高难度且易验证的任务
  • 在tau-bench等基准上显著提升训练稳定性和效率
  • 适合研究智能体强化学习与仿真环境设计的学者

在模拟环境中进行强化学习(RL)为提升基于语言的智能体提供了成本低、可扩展性强的方案。然而,以往工作受限于半自动化环境生成或任务难度不足,缺乏广度与深度。同时,模拟用户行为不稳定及环境异质性进一步加剧了智能体强化学习的挑战。本文提出:(1)一种统一的自动化、可扩展的模拟环境生成流程,支持高难度但易于验证的任务;(2)一种环境级强化学习算法,不仅能有效缓解用户不稳定性,还能在环境层面进行优势估计,从而提升训练效率与稳定性。在tau-bench、tau2-Bench和VitaBench等多个智能体基准上的综合评估验证了方法的有效性。深入分析还表明其具备出色的跨域泛化能力。

原文摘要 · Abstract (English)

Conducting reinforcement learning (RL) in simulated environments offers a cost-effective and highly scalable way to enhance language-based agents. However, previous work has been limited to semi-automated environment synthesis or tasks lacking sufficient difficulty, offering little breadth or depth. In addition, the instability of simulated users integrated into these environments, along with the heterogeneity across simulated environments, poses further challenges for agentic RL. In this work, we propose: (1) a unified pipeline for automated and scalable synthesis of simulated environments associated with high-difficulty but easily verifiable tasks; and (2) an environment level RL algorithm that not only effectively mitigates user instability but also performs advantage estimation at the environment level, thereby improving training efficiency and stability. Comprehensive evaluations on agentic benchmarks, including tau-bench, tau2-Bench, and VitaBench, validate the effectiveness of our proposed method. Further in-depth analyses underscore its out-of-domain generalization.

强化学习智能体环境生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。