在高保真企业模拟环境中训练AI代理,显著提升其泛化能力。
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
- 构建了包含2500+实体的复杂企业支持模拟环境CoreCraft。
- 模型训练后任务通过率从25.37%提升至36.76%,并在多个外部基准上实现显著增益。
- 适合关注通用智能体训练与真实工作场景适配的研究者和开发者。
我们证明,在高保真强化学习环境中训练的AI代理能产生超越训练分布的泛化能力。本文引入CoreCraft,EnterpriseBench中的首个环境,由Surge AI开发的智能体强化学习环境套件。CoreCraft是一个完整运行的企业级客户支持组织仿真系统,包含超过2500个实体、14类实体类型及23种独特工具,用于衡量AI代理是否能完成真实工作中所需的多步骤、领域特定任务。前沿模型如GPT-5.2和Claude Opus 4.6在满足所有专家编写评分标准的情况下,任务完成率低于30%。使用该环境,我们对GLM 4.6采用群体相对策略优化(GRPO)与自适应截断进行训练。单次训练迭代后,模型在保留评估任务上的任务通过率从25.37%提升至36.76%。更重要的是,这些改进可迁移至分布外基准:在BFCL Parallel上提升+4.5%,在Tau2-Bench Retail上提升+7.4%,在Tool Decathlon(Pass@1)上提升+6.8%。我们认为三个环境特性与观察到的迁移表现一致:以任务为中心的世界构建,优化多样性和挑战性;专家编写的评分标准支持可靠奖励计算;反映真实职业流程的企业工作流。结果表明,环境质量、多样性与真实性是实现可泛化智能体能力的关键因素。
原文摘要 · Abstract (English)
We show that training AI agents on high-fidelity reinforcement learning environments produces capabilities that generalize beyond the training distribution. We introduce CoreCraft, the first environment in EnterpriseBench, Surge AI's suite of agentic RL environments. CoreCraft is a fully operational enterprise simulation of a customer support organization, comprising over 2,500 entities across 14 entity types with 23 unique tools, designed to measure whether AI agents can perform the multi-step, domain-specific work that real jobs demand. Frontier models such as GPT-5.2 and Claude Opus 4.6 solve fewer than 30% of tasks when all expert-authored rubric criteria must be satisfied. Using this environment, we train GLM 4.6 with Group Relative Policy Optimization (GRPO) and adaptive clipping. After a single epoch of training, the model improves from 25.37% to 36.76% task pass rate on held-out evaluation tasks. More importantly, these gains transfer to out-of-distribution benchmarks: +4.5% on BFCL Parallel, +7.4% on Tau2-Bench Retail, and +6.8% on Tool Decathlon (Pass@1). We believe three environment properties are consistent with the observed transfer: task-centric world building that optimizes for diverse, challenging tasks; expert-authored rubrics enabling reliable reward computation; and enterprise workflows that reflect realistic professional patterns. Our results suggest that environment quality, diversity, and realism are key factors enabling generalizable agent capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。