首个评估大模型物理工具理解能力的基准,揭示现有模型严重不足。
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- 构建三阶视觉问答数据集,测试工具识别、原理理解与创造能力。
- 32个模型普遍表现不佳,工具理解能力远低于人类水平。
- 适合研究具身智能、多模态模型评估的学者和开发者参考。
使用、理解并创造工具是人类智能的核心特征,使我们能够与物理世界进行复杂交互。任何通用智能体若要实现真正多功能性,也必须掌握这些基本技能。尽管现代多模态大语言模型(MLLMs)利用其丰富的常识知识在具身AI及下游视觉-语言-动作(VLA)模型中实现高层规划,但其对物理工具的真实理解程度仍缺乏量化评估。为填补这一空白,我们提出PhysToolBench,首个专注于评估MLLMs对物理工具理解能力的基准。该基准采用超过1000组图像-文本对构成的视觉问答(VQA)数据集,涵盖三个难度层级:(1)工具识别:要求识别工具的主要功能;(2)工具理解:测试模型对工具工作原理的认知;(3)工具创造:挑战模型在缺乏常规工具时,利用周围物体构造新工具的能力。我们对32个MLLMs——包括专有模型、开源模型、具身专用模型及VLA基础模型——进行了全面评估,发现其工具理解能力存在显著缺陷。此外,我们提供深入分析并提出初步解决方案。代码与数据集已公开。
原文摘要 · Abstract (English)
The ability to use, understand, and create tools is a hallmark of human intelligence, enabling sophisticated interaction with the physical world. For any general-purpose intelligent agent to achieve true versatility, it must also master these fundamental skills. While modern Multimodal Large Language Models (MLLMs) leverage their extensive common knowledge for high-level planning in embodied AI and in downstream Vision-Language-Action (VLA) models, the extent of their true understanding of physical tools remains unquantified. To bridge this gap, we present PhysToolBench, the first benchmark dedicated to evaluating the comprehension of physical tools by MLLMs. Our benchmark is structured as a Visual Question Answering (VQA) dataset comprising over 1,000 image-text pairs. It assesses capabilities across three distinct difficulty levels: (1) Tool Recognition: Requiring the recognition of a tool's primary function. (2) Tool Understanding: Testing the ability to grasp the underlying principles of a tool's operation. (3) Tool Creation: Challenging the model to fashion a new tool from surrounding objects when conventional options are unavailable. Our comprehensive evaluation of 32 MLLMs-spanning proprietary, open-source, specialized embodied, and backbones in VLAs-reveals a significant deficiency in tool understanding. Furthermore, we provide an in-depth analysis and propose preliminary solutions. Code and dataset are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。