arXiv:2606.10803cs.CLcs.AI2026-06

首个物理工具使用评测基准,揭示大模型在真实场景中识物与规划能力的严重不足。

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

论文配图:Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use
图 1 · 摘自论文原文
  • 构建2678种真实工具的评测集,测试模型识物与任务规划能力。
  • 最强模型仅识别58.7%工具,仅完成21.0%任务,暴露感知与推理双重短板。
  • 揭示具身智能发展关键瓶颈:缺乏工具功能常识,适合研究机器人认知的学者参考。

多模态大语言模型(MLLMs)在使用数字API方面表现优异,并日益作为具身智能的“大脑”,指导机器人与物理世界交互。其中,物理工具的使用是核心能力,直接决定其在现实任务中的辅助价值。然而,当前对MLLMs在物理工具使用方面的表现仍缺乏系统评估。为此,我们提出PhysTool-Bench——首个针对物理工具使用的评测基准,用于评估模型理解真实场景、识别物理工具并规划使用顺序的能力。该基准包含2,510个问题,覆盖2,678种真实工具,涵盖制造、电气、农业和医疗等多个领域。评测从两个维度展开:1)识别场景中所有存在的工具;2)根据指令与视觉上下文规划工具选择与使用顺序。在13个主流MLLMs上测试发现,即使最强模型Gemini-3.1-Pro也仅能识别58.7%的工具,仅21.0%的任务可完整执行。分析表明存在双重缺陷:模型难以在真实场景中感知工具,且在规划阶段性能大幅下降,反映出其缺乏将感知工具映射到任务语义的功能常识,暴露出具身智能发展的关键瓶颈。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the "brain" of embodied AI, instructing robots to interact with the physical world. In such embodied settings, a central capability is the use of physical tools, which underpins MLLMs' ability to assist humans in real-world tasks. Despite the importance, MLLMs' proficiency in physical tool use remains largely unexplored. To address this gap, we introduce PhysTool-Bench, the first physical tool-use benchmark designed to evaluate MLLMs' ability to comprehend real-world scenarios, identify physical tools, and plan their use. PhysTool-Bench comprises 2,510 queries over 2,678 real-world physical tools spanning diverse domains, including manufacturing, electrical work, agriculture, and healthcare. Concretely, models are evaluated along two primary dimensions: 1) recognizing all physical tools present in the scene, and 2) planning the tool selection and use sequence based on the instruction and visual context. Across 13 leading MLLMs, even the strongest model (Gemini-3.1-Pro) identifies only 58.7% of tools in a scene and completes merely 21.0% of queries end-to-end. Our analysis reveals a two-level deficit: MLLMs struggle to perceive tools in realistic scenes, and the much larger drop at the planning stage further indicates a lack of functional commonsense for mapping perceived tools onto task semantics, pinpointing a critical bottleneck for the development of practical embodied AI.

具身智能工具使用多模态模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。