构建真实世界工具使用评估基准,考验多模态闭环纠错能力。
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents

- 设计闭环多模态验证机制,要求智能体执行工具后自检修正。
- 覆盖100个可执行任务,人类成功率94.0%,最强模型仅32.0%。
- 适合评估下一代多模态工具型智能体,尤其关注自纠错能力。
工具使用智能体正被期待应用于真实专业流程中,需处理多模态输入、协调外部工具、检查中间成果并修正行为以产出最终结果。现有基准往往孤立评估工具使用、计算机操作和多模态推理,难以反映真实世界端到端的全模态工具使用。为此,我们提出MM-ToolBench,一个面向任务导向的全模态工具使用基准与评估框架。该基准包含来自客户客服与智能创作两大任务族的100个可执行任务,覆盖20个子类别,由27个MCP服务器支持324种工具。核心设计为闭环多模态验证:智能体需执行工具、检查生成或转换的成果,并在未满足任务要求时自我修正。为实现可扩展且可验证的评估,MM-ToolBench结合MCP执行、任务特定的接地评估器,以及半自动化场景发现、任务实例化、评估器合成与人工审核流程。对15个主流代理模型的实验显示,该基准仍极具挑战性:常被视为最强编码代理的Claude Opus 4.6仅达32.0%任务成功率,远低于94.0%的人类基准。我们期望MM-ToolBench能成为推动下一代全模态工具型智能体发展的实用基础。
原文摘要 · Abstract (English)
Tool-using agents are increasingly expected to operate across realistic professional workflows, where they must interpret multimodal inputs, coordinate external tools, inspect intermediate artifacts, and revise their actions before producing a final result. Existing benchmarks, however, often evaluate tool use, computer use, and multimodal reasoning in isolation, leaving a gap between benchmark settings and end-to-end omni-modal tool use in the real world. To address this gap, we introduce MM-ToolBench, a benchmark and evaluation harness for task-oriented omni-modal tool use. MM-ToolBench contains 100 executable tasks from two macro task families, Customer Service and Intelligent Creation, covering 20 subcategory slices and supported by 27 MCP servers with 324 tools. The central design of MM-ToolBench is closed-loop multimodal verification: agents must execute tools, inspect rendered or transformed artifacts, and self-correct when outputs fail task-specific requirements. To make such evaluation scalable and verifiable, MM-ToolBench couples MCP-based execution with task-specific grounded evaluators and a semi-automated construction pipeline for scenario discovery, task instantiation, evaluator synthesis, and human audit. Experiments on 15 contemporary agentic models show that MM-ToolBench remains highly challenging: Claude Opus 4.6, commonly regarded as one of the strongest coding-agent models, achieves only 32.0% task success, far below the 94.0% human benchmark. We envision MM-ToolBench as a practical foundation for evaluating and advancing next-generation omni-modal tool-using agents through closed-loop multimodal verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。