arXiv:2603.15030cs.AI2026-03被引 12

评测多模态模型用工具完成复杂视觉任务的能力,发现当前模型严重依赖熟悉工具。

VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining

  • 设计32种基于OpenCV的视觉工具组合,模拟真实视觉处理流程
  • 19个模型测试显示顶尖模型仅达51%准确率,多工具协作仍难
  • 适合研究多模态智能体、视觉推理与工具链泛化能力的学者

近期进展将多模态大语言模型(MLLMs)从标准视觉问答拓展至利用外部工具完成高级视觉任务。然而,精确执行并有效组合多种工具以应对复杂任务仍是主要瓶颈。现有基准受限于工具集稀疏和使用路径简单,难以捕捉复杂多样的工具交互,无法在真实场景下评估模型性能。为此,我们提出VisualToolChain-Bench(VTC-Bench),一个全面评估MLLM工具使用能力的基准。框架包含32种基于OpenCV的多样化视觉操作,支持广泛组合,可严格评估多工具组合与长时序多步计划执行能力。为实现精准评估,我们提供680个经筛选的问题,涵盖九类认知层级,并附有真实执行轨迹。对19个领先MLLM的实验表明,当前模型在视觉智能体能力上存在显著局限:难以适应多样工具集,对未见过的操作泛化能力差,顶级模型Gemini-3.0-Pro仅达51%准确率。此外,面对复杂任务时,模型难以制定高效执行计划,过度依赖少量熟悉函数而非最优工具选择。VTC-Bench揭示了这些根本挑战,建立了指导更通用视觉智能体发展的严格基准。

原文摘要 · Abstract (English)

Recent advancements extend Multimodal Large Language Models (MLLMs) beyond standard visual question answering to utilizing external tools for advanced visual tasks. Despite this progress, precisely executing and effectively composing diverse tools for complex tasks remain persistent bottleneck. Constrained by sparse tool-sets and simple tool-use trajectories, existing benchmarks fail to capture complex and diverse tool interactions, falling short in evaluating model performance under practical, real-world conditions. To bridge this gap, we introduce VisualToolChain-Bench(VTC-Bench), a comprehensive benchmark designed to evaluate tool-use proficiency in MLLMs. To align with realistic computer vision pipelines, our framework features 32 diverse OpenCV-based visual operations. This rich tool-set enables extensive combinations, allowing VTC-Bench to rigorously assess multi-tool composition and long-horizon, multi-step plan execution. For precise evaluation, we provide 680 curated problems structured across a nine-category cognitive hierarchy, each with ground-truth execution trajectories. Extensive experiments on 19 leading MLLMs reveal critical limitations in current models' visual agentic capabilities. Specifically, models struggle to adapt to diverse tool-sets and generalize to unseen operations, with the leading model Gemini-3.0-Pro only achieving 51% on our benchmark. Furthermore, multi-tool composition remains a persistent challenge. When facing complex tasks, models struggle to formulate efficient execution plans, relying heavily on a narrow, suboptimal subset of familiar functions rather than selecting the optimal tools. By identifying these fundamental challenges, VTC-Bench establishes a rigorous baseline to guide the development of more generalized visual agentic models.

多模态模型工具链视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。