构建物理感知分类体系,评估模型对物体属性的深层理解能力。
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
- 设计分层场景推理框架,涵盖物体、空间与属性三类认知层级。
- 在5802张图像上构建28033个问题,测试模型对材料、功能等属性的推理表现。
- 发现主流模型在多步属性推理中性能下降10%-20%,提示其依赖模式匹配而非真实理解。
我们提出感知分类(Perceptual Taxonomy),一种结构化的场景理解流程:先识别物体及其空间关系,再推断与任务相关的属性,如材质、可操作性、功能和物理特征,以支持目标导向的推理。尽管这种推理是人类认知的基础,但现有视觉语言基准仅关注表面识别或图文对齐,缺乏对此类能力的全面评估。为此,我们构建了首个基于物理实体的视觉推理基准。标注了3173个物体的四类属性家族,覆盖84种细粒度属性,并基于这些标注构建了一个多项选择题基准,包含5802张合成与真实图像。该基准涵盖28033个基于模板的问题,分为四类:物体描述、空间推理、属性匹配和分类推理,另含50个专家设计的问题,用于全面评估模型在感知分类推理中的表现。实验表明,领先视觉语言模型在识别任务中表现良好,但在属性驱动的问题上性能下降10%至20%,尤其在涉及结构化属性的多步推理中更为明显。这揭示了当前模型在结构化视觉理解上的持续差距,以及其过度依赖模式匹配的局限性。此外,提供模拟场景中的上下文推理示例能显著提升模型在真实世界和专家问题上的表现,证明了感知分类引导提示的有效性。
原文摘要 · Abstract (English)
We propose Perceptual Taxonomy, a structured process of scene understanding that first recognizes objects and their spatial configurations, then infers task-relevant properties such as material, affordance, function, and physical attributes to support goal-directed reasoning. While this form of reasoning is fundamental to human cognition, current vision-language benchmarks lack comprehensive evaluation of this ability and instead focus on surface-level recognition or image-text alignment. To address this gap, we introduce Perceptual Taxonomy, a benchmark for physically grounded visual reasoning. We annotate 3173 objects with four property families covering 84 fine-grained attributes. Using these annotations, we construct a multiple-choice question benchmark with 5802 images across both synthetic and real domains. The benchmark contains 28033 template-based questions spanning four types (object description, spatial reasoning, property matching, and taxonomy reasoning), along with 50 expert-crafted questions designed to evaluate models across the full spectrum of perceptual taxonomy reasoning. Experimental results show that leading vision-language models perform well on recognition tasks but degrade by 10 to 20 percent on property-driven questions, especially those requiring multi-step reasoning over structured attributes. These findings highlight a persistent gap in structured visual understanding and the limitations of current models that rely heavily on pattern matching. We also show that providing in-context reasoning examples from simulated scenes improves performance on real-world and expert-curated questions, demonstrating the effectiveness of perceptual-taxonomy-guided prompting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。