测试视觉语言模型的视角转换能力,发现多数模型存在自我中心偏差。
Egocentric Bias in Vision-Language Models
- 设计新基准FlipSet,通过旋转2D字符测试模型视角转换。
- 超75%错误重复摄像头视角,显示严重自我中心倾向。
- 模型能单独完成心理旋转但无法整合社会认知与空间操作。
视觉视角转换——推断世界从他人视角呈现的方式——是社会认知的基础。我们提出FlipSet,一个针对二级视觉视角转换(L2 VPT)的诊断基准。该任务要求模拟从另一智能体视角对2D字符字符串进行180度旋转,将空间变换与3D场景复杂性分离。评估103个视觉语言模型(VLMs)发现系统性自我中心偏差:绝大多数表现低于随机水平,约四分之三的错误重现了摄像机视角。控制实验揭示组合缺陷——模型在孤立状态下能达到高理论心智准确率和高于随机的心理旋转能力,但在需要整合时却出现灾难性失败。这种分离表明,当前的VLM缺乏将社交意识与空间操作结合的机制,暗示其基于模型的空间推理存在根本局限。FlipSet为多模态系统中视角理解能力的诊断提供了认知上合理的基础测试平台。
原文摘要 · Abstract (English)
Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in vision-language models. The task requires simulating 180-degree rotations of 2D character strings from another agent's perspective, isolating spatial transformation from 3D scene complexity. Evaluating 103 VLMs reveals systematic egocentric bias: the vast majority perform below chance, with roughly three-quarters of errors reproducing the camera viewpoint. Control experiments expose a compositional deficit--models achieve high theory-of-mind accuracy and above-chance mental rotation in isolation, yet fail catastrophically when integration is required. This dissociation indicates that current VLMs lack the mechanisms needed to bind social awareness to spatial operations, suggesting fundamental limitations in model-based spatial reasoning. FlipSet provides a cognitively grounded testbed for diagnosing perspective-taking capabilities in multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。