评测大模型智能体在真实复杂环境中的实战能力
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

- 构建真实场景任务,测试智能体自主探索与工具发现能力
- 15个主流模型在新基准上表现普遍不佳,顶尖模型也难胜任
- 适合关注智能体落地应用的开发者和研究者参考
语言智能体(LLM agents)发展迅速,正逐步投入实际应用。然而,现有评估多基于简化、理想化环境,依赖预设工具接口,忽略关键环节,假设输入干净且完整。这导致评估结果低估了真实部署中的挑战——不确定性与噪声普遍存在,智能体需主动探索以发现新工具。为此,我们提出AgentGym2,一个基于真实端到端工作需求的任务框架。它不仅评估推理与规划能力,还考察智能体执行全流程操作、通过探索发现工具、组合工具应对未见任务,以及在噪声和信息不全情况下的鲁棒性。对15个专有及开源模型的实验表明,即使领先模型如Gemini和GPT-5在AgentGym2上也表现不佳,揭示当前智能体能力与真实应用需求间存在显著差距。
原文摘要 · Abstract (English)
Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic evaluations. However, most existing benchmarks evaluate agents in simplified, idealized settings. They typically rely on pre-packaged tool interfaces, overlook critical steps, and assume inputs are clean and fully specified. Consequently, they understate the difficulty of real deployments, where uncertainty and noise are ubiquitous and agents must proactively explore the environment to uncover new tools. To bridge this gap, we present AgentGym2, a new evaluation framework with task instances grounded in real-world end-to-end working demands. Beyond reasoning and planning, it measures agents' ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information. Experiments on 15 proprietary and open-source models show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2, revealing a substantial gap between the capability of current agents and the demands of real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。