测试大模型是否真懂物体操作的物理前提条件。
PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies?
- 构建新基准PAC Bench,评估视觉语言模型对物体属性、功能和约束的理解。
- 3万+标注数据揭示当前模型在物理常识上存在明显短板。
- 适合机器人、AI推理方向研究者参考,推动更可靠的智能体开发。
视觉语言模型(VLMs)正日益成为通用机器人操作的核心,支持物理推理、策略生成和故障检测等任务。然而,其在这些高阶应用中的表现通常依赖于对低阶物理前提的深层理解,而这一能力尚未得到充分验证。机器人要可靠执行动作,需掌握物体固有属性(如材质、重量)、动作可操作性(如可抓取、可堆叠)以及物理约束(如稳定性、可达性或状态,如是否关闭)。尽管VLM广泛用于操作任务,我们指出现成模型可能缺乏这种细粒度的物理基础理解,因训练中常忽略此类前提。为此,我们提出PAC Bench,一个系统评估VLM在任务可执行性视角下对核心属性(Properties)、可操作性(Affordances)与约束(Constraints)理解的综合性基准。该基准包含超过3万条标注,涵盖673张真实世界图像(115个物体类别,15种属性类型,每类定义1至3个可操作性),100个真实人形视角场景,以及跨四个任务的120个独特模拟约束场景。评估显示,当前VLM在掌握基本物理概念方面存在显著差距,暴露出其在可靠机器人操作中的适用局限,并指明关键研究方向。PAC Bench亦可作为标准化基准,用于严谨评估VLM的物理推理能力,指导开发更具鲁棒性的物理感知模型。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly pivotal for generalist robot manipulation, enabling tasks such as physical reasoning, policy generation, and failure detection. However, their proficiency in these high-level applications often assumes a deep understanding of low-level physical prerequisites, a capability that remains largely unverified. For robots to perform actions reliably, they must comprehend intrinsic object properties (e.g., material, weight), action affordances (e.g., graspable, stackable), and physical constraints (e.g., stability, reachability, or an object's state, such as being closed). Despite the widespread use of VLMs in manipulation tasks, we argue that off-the-shelf models may lack this granular, physically grounded understanding, as such prerequisites are often overlooked during training. To address this critical gap, we introduce PAC Bench, a comprehensive benchmark designed to systematically evaluate VLMs on their understanding of core Properties, Affordances, and Constraints (PAC) from a task executability perspective. PAC Bench features a diverse dataset with over 30,000 annotations, comprising 673 real-world images (115 object classes, 15 property types, and 1 to 3 affordances defined per class), 100 real-world humanoid-view scenarios, and 120 unique simulated constraint scenarios across four tasks. Our evaluations reveal significant gaps in the ability of current VLMs to grasp fundamental physical concepts, highlighting limitations in their suitability for reliable robot manipulation and pointing to key areas for targeted research. PAC Bench also serves as a standardized benchmark for rigorously evaluating physical reasoning in VLMs and guiding the development of more robust, physically grounded models for robotic applications. Project Page: https://pacbench.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。