arXiv:2509.13311cs.CL2025-09ACL被引 49

通过扩展环境提升智能体通用能力,实现更精准的函数调用。

Towards General Agentic Intelligence via Environment Scaling

  • 构建可自动扩展的多样化仿真环境,系统性覆盖函数调用场景。
  • 在tau-bench等基准上,模型函数调用准确率显著提升。
  • 适合研究通用智能体与实际应用部署的开发者参考。

高级智能体智能是将大语言模型应用于实际、真实世界应用的前提。多样化的现实世界API要求智能体具备精确、鲁棒的函数调用能力,这需要智能体通过在多样化环境中的交互来发展这些能力。函数调用能力的广度与训练环境的多样性密切相关。本文通过环境扩展作为迈向通用智能体智能的一步,提出两个核心挑战:(i) 如何有原则地扩展环境,(ii) 如何有效从与这些环境交互的经验中训练智能体能力。为此,我们设计了一个可扩展框架,能自动构建完全仿真的异构环境,系统性拓宽函数调用场景空间。我们进一步采用两阶段智能体微调策略:首先赋予智能体基础智能体能力,再针对特定领域进行专业化训练。在agentic基准tau-bench、tau2-Bench和ACEBench上的大量实验表明,所训练模型AgentScaler显著提升了模型的函数调用能力。

原文摘要 · Abstract (English)

Advanced agentic intelligence is a prerequisite for deploying Large Language Models in practical, real-world applications. Diverse real-world APIs demand precise, robust function-calling intelligence, which needs agents to develop these capabilities through interaction in varied environments. The breadth of function-calling competence is closely tied to the diversity of environments in which agents are trained. In this work, we scale up environments as a step towards advancing general agentic intelligence. This gives rise to two central challenges: (i) how to scale environments in a principled manner, and (ii) how to effectively train agentic capabilities from experiences derived through interactions with these environments. To address these, we design a scalable framework that automatically constructs heterogeneous environments that are fully simulated, systematically broadening the space of function-calling scenarios. We further adapt a two-phase agent fine-tuning strategy: first endowing agents with fundamental agentic capabilities, then specializing them for domain-specific contexts. Extensive experiments on agentic benchmarks, tau-bench, tau2-Bench, and ACEBench, demonstrate that our trained model, AgentScaler, significantly enhances the function-calling capability of models.

智能体函数调用环境扩展大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。