arXiv:2506.21458cs.AIcs.CL2025-06被引 76

让视觉语言模型像人一样从少视角推断完整空间布局

MindCube: Spatial Mental Modeling from Limited Views

  • 用认知地图+推理链联合训练,构建内部空间表征
  • 准确率从37.8%提升至61.3%,突破现有模型极限
  • 适合研究空间理解、具身智能与多视角推理的学者

视觉语言模型能否像人类一样,仅凭少数视角就构想出完整场景?我们构建了包含21,154个问题和3,268张图像的MindCube基准,揭示了现有模型在空间心理建模上的巨大差距——表现接近随机。通过该基准,我们系统评估模型在位置表征(认知制图)、方向理解(视角转换)和动态模拟(假设性运动推理)方面的表现。提出三种改进策略:引入未见中间视角、自然语言推理链和认知地图。最优方案为“先制图后推理”协同训练,使准确率从37.8%提升至57.8%(+20.0%)。加入强化学习后进一步达61.3%(+23.5%)。核心发现:主动构建并利用结构化内部空间表征,结合灵活推理,能显著增强对不可见空间的理解能力。

原文摘要 · Abstract (English)

Can Vision-Language Models (VLMs) imagine the full scene from just a few views, like humans do? Humans form spatial mental models naturally, internal representations of unseen space, to reason about layout, perspective, and motion. Our MindCube benchmark with 21,154 questions across 3,268 images exposes this critical gap, where existing VLMs exhibit near-random performance. Using MindCube, we systematically evaluate how well VLMs build robust spatial mental models through representing positions (cognitive mapping), orientations (perspective-taking), and dynamics (mental simulation for "what-if" movements). We then explore three approaches to help approximate spatial mental models in VLMs, focusing on incorporating unseen intermediate views, natural language reasoning chains, and cognitive maps. The significant improvement comes from a synergistic approach, "map-then-reason", that jointly trains the model to first generate a cognitive map and then reason upon it. By training models to reason over these internal maps, we boosted accuracy from 37.8% to 57.8% (+20.0%). Adding reinforcement learning pushed performance even further to 61.3% (+23.5%). Our key insight is that such scaffolding of spatial mental models, actively constructing and utilizing internal structured spatial representations with flexible reasoning processes, significantly improves understanding of unobservable space.

空间建模视觉语言模型认知地图多视角推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。