对比视觉语言模型与人类对几何概念的理解差异
Decoupling the components of geometric understanding in Vision Language Models
- 用认知科学方法分离几何理解与推理等能力
- 模型在多数几何任务上逊于受过教育的成人和无教育背景的原住民
- 模型在需要心理旋转的任务中表现更差,体现理解脆性
几何理解高度依赖视觉。本文评估当前最先进的视觉语言模型(VLMs)是否能理解简单的几何概念。采用认知科学中的范式,将简单几何的视觉理解与其他常被混淆的能力(如推理、世界知识)分离。比较模型表现与美国成人的表现,以及亚马逊原住民群体中无正式教育经历的成人的表现。结果发现,尽管部分几何概念上模型表现尚可,但总体上始终低于两组人类;且模型的几何理解更为脆弱,在需要心理旋转的任务中表现不佳。该研究揭示了人类与机器在几何理解起源上的差异——前者可能源于书面材料与物理互动的结合,而后者则缺乏这种多源经验,为理解二者差异迈出一小步。
原文摘要 · Abstract (English)
Understanding geometry relies heavily on vision. In this work, we evaluate whether state-of-the-art vision language models (VLMs) can understand simple geometric concepts. We use a paradigm from cognitive science that isolates visual understanding of simple geometry from the many other capabilities it is often conflated with such as reasoning and world knowledge. We compare model performance with human adults from the USA, as well as with prior research on human adults without formal education from an Amazonian indigenous group. We find that VLMs consistently underperform both groups of human adults, although they succeed with some concepts more than others. We also find that VLM geometric understanding is more brittle than human understanding, and is not robust when tasks require mental rotation. This work highlights interesting differences in the origin of geometric understanding in humans and machines -- e.g. from printed materials used in formal education vs. interactions with the physical world or a combination of the two -- and a small step toward understanding these differences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。