arXiv:2602.06820cs.AI2026-02被引 12

从零构建可交互的多样化环境,提升智能体泛化能力

ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training

  • 通过程序化测试保证环境可靠性,构建可验证任务
  • 在τ²-Bench和VitaBench上实现显著性能提升
  • 证明环境多样性规模直接影响模型泛化效果

训练能够适应多种场景的通用智能体需要交互式环境以支持自主探索。然而,当前交互式环境仍极度稀缺,现有合成方法在环境多样性和可扩展性方面存在严重局限。为此,我们提出ScaleEnv框架,从零开始构建完全可交互的环境与可验证的任务。ScaleEnv通过程序化测试确保环境可靠性,并利用工具依赖图扩展和可执行动作验证,保障任务的完整性和可解性。通过在ScaleEnv中进行探索学习,智能体在未见过的多轮工具使用基准(如τ²-Bench和VitaBench)上表现显著提升,展现出强大的泛化能力。此外,我们研究了领域数量增加与模型泛化性能的关系,提供了实证证据:扩大环境多样性对鲁棒智能体学习至关重要。

原文摘要 · Abstract (English)

Training generalist agents capable of adapting to diverse scenarios requires interactive environments for self-exploration. However, interactive environments remain critically scarce, and existing synthesis methods suffer from significant limitations regarding environmental diversity and scalability. To address these challenges, we introduce ScaleEnv, a framework that constructs fully interactive environments and verifiable tasks entirely from scratch. Specifically, ScaleEnv ensures environment reliability through procedural testing, and guarantees task completeness and solvability via tool dependency graph expansion and executable action verification. By enabling agents to learn through exploration within ScaleEnv, we demonstrate significant performance improvements on unseen, multi-turn tool-use benchmarks such as $τ^2$-Bench and VitaBench, highlighting strong generalization capabilities. Furthermore, we investigate the relationship between increasing number of domains and model generalization performance, providing empirical evidence that scaling environmental diversity is critical for robust agent learning.

智能体训练环境合成泛化能力工具使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。