测试视觉语言模型的视角理解能力,发现其识别物体易但理解空间视角难。
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
- 用积木人与物体组合设计144个可控视觉任务
- 模型在视角理解上表现远差于场景认知
- 适合研究视觉推理与模型认知边界的学者
我们通过受控场景测试视觉语言模型(VLMs)的视觉视角理解能力,场景中单一人形积木人与单一物体组合,系统性改变物体位置与积木人朝向,并使用鸟瞰与表面视角构建144个独特任务。每个任务配以7个诊断问题,评估三个层次的视觉认知:场景理解、空间推理和视角理解。测试包括Gemini Robotics-ER 1.5、Llama-3.2-11B-Vision-Instruct、Claude Sonnet、GPT-4和Qwen3等高性能模型。结果显示,模型在场景理解上表现良好,但在空间推理上性能下降,视角理解更显著恶化。分析表明,当前模型在表层物体识别与深层空间及视角推理之间存在差距,提示未来需引入显式几何表示与定制化训练协议。
原文摘要 · Abstract (English)
We investigate the ability of Vision Language Models (VLMs) to perform visual perspective taking using a new set of visual tasks inspired by established human tests. Our approach leverages carefully controlled scenes in which a single humanoid minifigure is paired with a single object. By systematically varying spatial configurations -- such as object position relative to the minifigure and the minifigure's orientation -- and using both bird's-eye and surface-level views, we created 144 unique visual tasks. Each task is paired with a series of 7 diagnostic questions designed to assess three levels of visual cognition: scene understanding, spatial reasoning, and visual perspective taking. We evaluate several high-performing models, including Gemini Robotics-ER 1.5, Llama-3.2-11B-Vision-Instruct, and variants of Claude Sonnet, GPT-4, and Qwen3, and find that while they excel at scene understanding, performance declines markedly on spatial reasoning and deteriorates further on perspective taking. Our analysis suggests a gap between surface-level object recognition and the deeper spatial and perspective reasoning required for complex visual tasks, pointing to the need for integrating explicit geometric representations and tailored training protocols in future VLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。