构建视觉物体属性常识推理新基准,揭示大模型在计数与比较任务中的显著短板。
A Study of Commonsense Reasoning over Visual Object Properties
- 设计三类图像、三层次推理、四属性维度的系统性评估框架
- 12个顶尖模型零样本测试,计数准确率不足40%,对比任务差20%人类水平
- 特别在真实照片、功能属性和高数量场景下表现薄弱,适合研究视觉推理局限者
受人类分类启发,对物体属性(如物理特性与功能)的视觉推理需识别低层细节与高层抽象。现有视觉问答(VQA)研究虽涵盖大小等属性,但常混淆感知与推理,且缺乏推理层级与图像类别的代表性,难以判断视觉语言模型(VLMs)如何识别与推理图像对象。为此,我们提出系统性评估框架:包含三类代表性图像、三阶递增复杂度推理层级及四维物体属性,基于先前常识知识表征与推理研究。我们构建两个基准:OPTICS-CNT含360张图像与1,080个计数类多层级问题;OPTICS-CMP含2.1k个比较类问题。在12个先进VLM的零样本实验中,最佳模型计数准确率低于40%,比较准确率约70%,与人类仍有20%差距。模型尤其在摄影图像、反事实推理、物理与功能属性以及高数量场景中表现不佳。我们公开OPTICS基准数据与代码,以支持未来可扩展的基准方法、泛化标注规范与先进推理型VLM发展。
原文摘要 · Abstract (English)
Inspired by human categorization, visual reasoning about object properties, such as physical attributes and functions, involves identifying and recognizing low-level details and higher-level abstractions. While current visual question answering (VQA) studies consider multiple object properties, such as size, they typically blend perception and reasoning and lack representativeness with respect to reasoning levels and image categories, making it unclear whether and how vision-language models (VLMs) recognize and reason about depicted objects. To this end, we introduce a systematic evaluation framework comprising images of three representative types, three reasoning levels of increasing complexity, and four object property dimensions, informed by prior work on commonsense knowledge representation and reasoning. We develop a procedure to instantiate this framework in two VQA object-reasoning benchmarks: OPTICS-CNT, comprising 360 images paired with 1,080 multi-level, count-based questions, and OPTICS-CMP, comprising 2.1k comparison questions. Experiments with 12 state-of-the-art VLMs in zero-shot settings reveal significant limitations relative to humans, with the best-performing model achieving below 40% counting and 70% comparison accuracy. While newer reasoning models perform better, a 20% gap to human performance remains. VLMs struggle particularly with photographic images, counterfactual reasoning, physical and functional properties, and higher counts. We make the OPTICS benchmark data and code available to support future scalable benchmarking methods, generalized annotation guidelines, and advanced reasoning VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。