arXiv:2601.16378cs.CVcs.AI2026-01被引 1

通过仿人认知的视角标记,解决多模态模型的自我中心偏差问题。

Cognitively-Inspired Tokens Overcome Egocentric Bias in Multimodal Models

  • 引入视角标记,用身体关键点或抽象旋转表示方向。
  • 在合成与真实场景中提升视角转换任务准确率,支持非人类视角。
  • 轻量级设计,适配不同模型,推动更类人的空间推理。

多模态语言模型在语义视觉-语言任务上表现良好,但在需要采用其他智能体视觉视角的空间推理任务中表现不佳,反映出持续存在的自我中心偏差,引发对当前模型是否支持非自我中心推理的质疑。受人类空间认知启发,我们提出视角标记,即通过(1)具身身体关键点线索或(2)支持心理旋转的抽象表征来编码方向的专用嵌入。将这些标记集成到 LLaVA-1.5-13B 模型中,使其在二级视觉视角转换任务上达到较高性能。在合成数据集 Isle Bricks V2 与自然场景数据集 COCO、3DSRBench 上,视角标记均显著提升准确率,基于旋转的标记还能泛化至非人类参考智能体。表征分析显示,微调增强了基线模型中已存在的潜在方向敏感性,表明多模态语言模型中潜藏非自我中心推理的雏形,但缺乏合适的内部结构。总体而言,将基于认知的空间结构直接嵌入标记空间,为视角转换和更类人的空间推理提供了一种轻量、模型无关的机制。

原文摘要 · Abstract (English)

Multimodal language models (MLMs) perform well on semantic vision-language tasks but fail at spatial reasoning that requires adopting another agent's visual perspective. These errors reflect a persistent egocentric bias and raise questions about whether current models support allocentric reasoning. Inspired by human spatial cognition, we introduce perspective tokens, specialized embeddings that encode orientation through either (1) embodied body-keypoint cues or (2) abstract representations supporting mental rotation. Integrating these tokens into LLaVA-1.5-13B yields performance on level-2 visual perspective-taking tasks. Across synthetic and naturalistic benchmarks (Isle Bricks V2, COCO, 3DSRBench), perspective tokens improve accuracy, with rotation-based tokens generalizing to non-human reference agents. Representational analyses reveal that fine-tuning enhances latent orientation sensitivity already present in the base model, suggesting that MLMs contain precursors of allocentric reasoning but lack appropriate internal structure. Overall, embedding cognitively grounded spatial structure directly into token space provides a lightweight, model-agnostic mechanism for perspective-taking and more human-like spatial reasoning.

空间推理视角转换多模态认知启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。