arXiv:2604.15715cs.CLcs.AI2026-04被引 1

构建真实工具使用基准,评估智能体从单步操作到复杂流程的全流程能力。

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

论文配图:GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
图 1 · 摘自论文原文
  • 分层级设计:原子任务与开放流程并重,贴近真实工作场景。
  • 顶尖模型在流程任务中成功率仅14.39%,远低于原子任务。
  • 执行框架比模型能力更关键,反馈机制可显著提升表现。

通用智能体的发展需从执行简单指令转向完成复杂现实生产力流程。然而,现有工具使用基准仍与真实需求脱节,依赖生成查询、虚拟工具和有限系统协作。为此,我们提出GTA-2,一个涵盖原子工具使用与开放流程的分层基准。基于真实世界数据,它采用真实用户查询、已部署工具及多模态上下文。(i) GTA-Atomic继承自先前基准,评估短周期、封闭式任务的精度;(ii) GTA-Workflow引入长周期、开放式任务,实现端到端真实流程完成。为评估开放输出,我们提出基于递归检查点的评估机制,将目标分解为可验证子目标,统一评估模型能力与执行框架(即执行架构)。实验显示明显能力断崖:尽管前沿模型在原子任务中已难达标(低于50%),在流程任务中表现更差,顶级模型成功率仅为14.39%。进一步分析表明,检查点引导反馈能提升性能,而Manus与OpenClaw等先进框架显著增强流程完成率,凸显执行架构设计的重要性超越模型本身。该研究为构建可靠个人与专业助手提供指导。数据集与代码将在https://github.com/open-compass/GTA公开。

原文摘要 · Abstract (English)

The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows. However, current tool-use benchmarks remain misaligned with real-world requirements, relying on AI-generated queries, dummy tools, and limited system-level coordination. To address this, we propose GTA-2, a hierarchical benchmark for General Tool Agents (GTA) spanning atomic tool use and open-ended workflows. Built on real-world authenticity, it leverages real user queries, deployed tools, and multimodal contexts. (i) GTA-Atomic, inherited from our prior GTA benchmark, evaluates short-horizon, closed-ended tool-use precision. (ii) GTA-Workflow introduces long-horizon, open-ended tasks for realistic end-to-end completion. To evaluate open-ended deliverables, we propose a recursive checkpoint-based evaluation mechanism that decomposes objectives into verifiable sub-goals, enabling unified evaluation of both model capabilities and agent execution frameworks (i.e., execution harnesses). Experiments reveal a pronounced capability cliff: while frontier models already struggle on atomic tasks (below 50%), they largely fail on workflows, with top models achieving only 14.39% success. Further analysis shows that checkpoint-guided feedback improves performance, while advanced frameworks such as Manus and OpenClaw substantially enhance workflow completion, highlighting the importance of execution harness design beyond the underlying model capacity. These findings provide guidance for developing reliable personal and professional assistants. Dataset and code will be available at https://github.com/open-compass/GTA.

智能体评估工具使用工作流基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。