arXiv:2510.25726cs.CLcs.AI2025-10被引 65

构建复杂真实任务的评估基准,测试语言代理长程执行能力。

The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

  • 设计涵盖32个应用、604个工具的跨平台任务集
  • 平均需20步交互完成任务,顶尖模型成功率仅38.6%
  • 基于真实环境状态,适合评估真实世界长程任务表现

现实世界中的语言代理需在多样化应用间完成多步骤复杂任务,如协调日历与文件系统管理邮件,或监控数据库并按手册生成报告。现有基准多聚焦于狭窄领域或简化任务,缺乏多样性、真实性和长周期复杂性。为此,我们提出Tool Decathlon(Toolathlon)基准,包含32个软件应用和604个工具,覆盖从Google Calendar、Notion到WooCommerce、Kubernetes、BigQuery等日常及专业平台,多数基于自建或优化的Model Context Protocol(MCP)服务器。区别于以往仅保证功能真实但环境状态单一的方案,本基准引入真实初始状态,如含数十名学生的Canvas课程或真实财务表格。共设108项人工设计任务,平均需约20步交互完成,每项任务通过专用评估脚本严格验证。对主流模型的全面评估显示显著不足:最佳模型Claude-4.5-Sonnet成功率为38.6%,平均调用20.2次工具;顶级开源模型DeepSeek-V3.2-Exp为20.1%。该基准有望推动更强大语言代理的发展。

原文摘要 · Abstract (English)

Real-world language agents must handle complex, multi-step workflows across diverse Apps. For instance, an agent may manage emails by coordinating with calendars and file systems, or monitor a production database to detect anomalies and generate reports following an operating manual. However, existing language agent benchmarks often focus on narrow domains or simplified tasks that lack the diversity, realism, and long-horizon complexity required to evaluate agents' real-world performance. To address this gap, we introduce the Tool Decathlon (dubbed as Toolathlon), a benchmark for language agents offering diverse Apps and tools, realistic environment setup, and reliable execution-based evaluation. Toolathlon spans 32 software applications and 604 tools, ranging from everyday platforms such as Google Calendar and Notion to professional ones like WooCommerce, Kubernetes, and BigQuery. Most of the tools are based on a high-quality set of Model Context Protocol (MCP) servers that we may have revised or implemented ourselves. Unlike prior works, which primarily ensure functional realism but offer limited environment state diversity, we provide realistic initial environment states from real software, such as Canvas courses with dozens of students or real financial spreadsheets. This benchmark includes 108 manually sourced or crafted tasks in total, requiring interacting with multiple Apps over around 20 turns on average to complete. Each task is strictly verifiable through dedicated evaluation scripts. Comprehensive evaluation of SOTA models highlights their significant shortcomings: the best-performing model, Claude-4.5-Sonnet, achieves only a 38.6% success rate with 20.2 tool calling turns on average, while the top open-weights model DeepSeek-V3.2-Exp reaches 20.1%. We expect Toolathlon to drive the development of more capable language agents for real-world, long-horizon task execution.

语言代理任务评估长程执行多工具协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。