测试多模态模型在视觉空间视角转换中的能力,发现其严重不足。
Visuospatial Perspective Taking in Multimodal Language Models
- 用人类研究任务改编评估模型视角转换能力
- 模型在第二层级视角转换中表现明显落后
- 对协作场景下模型应用有重要警示意义
随着多模态语言模型(MLMs)在社交与协作场景中的广泛应用,评估其视角转换能力变得至关重要。现有基准主要依赖文本情景或静态场景理解,导致视觉空间视角转换(VPT)研究不足。我们借鉴人类研究中的两项任务:导演任务(Director Task),用于评估参照性交流中的视角转换;旋转图形任务(Rotating Figure Task),考察角度差异下的视角转换能力。结果显示,各类MLMs在第二层级的视觉空间视角转换中存在显著缺陷,需抑制自身视角以采纳他人视角。这揭示了当前MLMs在表征和推理他人视角方面存在关键局限,对其在协作场景中的应用带来深远影响。
原文摘要 · Abstract (English)
As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks largely rely on text-based vignettes or static scene understanding, leaving visuospatial perspective-taking (VPT) underexplored. We adapt two evaluation tasks from human studies: the Director Task, assessing VPT in a referential communication paradigm, and the Rotating Figure Task, probing perspective-taking across angular disparities. Across tasks, MLMs show pronounced deficits in Level 2 VPT, which requires inhibiting one's own perspective to adopt another's. These results expose critical limitations in current MLMs' ability to represent and reason about alternative perspectives, with implications for their use in collaborative contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。