arXiv:2603.27183cs.CV2026-03

让大模型通过对话整合不同视角信息,构建共享空间认知。

Communicating about Space: Language-Mediated Spatial Integration Across Partial Views

  • 用对话协同整合多视角观察,构建统一空间模型
  • 顶尖模型仅72%准确率,全局地图构建基本失败
  • 人类对话渐趋精准,模型却难收敛,缺乏共享心智

人类通过交流部分、视角依赖的观察来建立共享的空间理解。我们探究多模态大语言模型(MLLMs)是否具备类似能力,即通过对话对齐不同第一人称视角,形成一致的客观空间认知。为此,我们提出COSMIC基准,用于协作式空间通信研究。该设置中,两个静态的MLLM代理从不同视角观察同一3D室内环境,并通过自然语言交流解答空间问题。COSMIC包含899个多样场景和1250组问答对,涵盖五类任务。实验发现:MLLMs在识别跨视角共现锚点物体上表现最佳,关系推理能力较弱,而全局一致性地图构建几乎完全失败,接近随机水平,即使前沿模型也如此。此外,思维链能力有助于锚点定位,但不足以支撑高层空间对话。为理解模型行为,我们收集了250组真人对话。人类整体准确率达95%,最佳模型Gemini-3-Pro-Thinking仅达72%,仍有巨大提升空间。人类对话随理解同步而愈发精确,而模型持续探索却不收敛,表明其难以建立并维持稳定的共享心智模型。代码与数据已公开于https://github.com/ankursikarwar/Cosmic。

原文摘要 · Abstract (English)

Humans build shared spatial understanding by communicating partial, viewpoint-dependent observations. We ask whether Multimodal Large Language Models (MLLMs) can do the same, aligning distinct egocentric views through dialogue to form a coherent, allocentric mental model of a shared environment. To study this systematically, we introduce COSMIC, a benchmark for Collaborative Spatial Communication. In this setting, two static MLLM agents observe a 3D indoor environment from different viewpoints and exchange natural-language messages to solve spatial queries. COSMIC contains 899 diverse scenes and 1250 question-answer pairs spanning five tasks. We find a capability hierarchy, MLLMs are most reliable at identifying shared anchor objects across views, perform worse on relational reasoning, and largely fail at building globally consistent maps, performing near chance, even for frontier models. Moreover, we find thinking capability yields gains in anchor grounding, but is insufficient for higher-level spatial communication. To contextualize model behavior, we collect 250 human-human dialogues. Humans achieve 95% aggregate accuracy, while the best model, Gemini-3-Pro-Thinking, reaches 72%, leaving substantial room for improvement. Moreover, human conversations grow more precise as partners align on a shared spatial understanding, whereas MLLMs keep exploring without converging, suggesting limited capacity to form and sustain a robust shared mental model throughout the dialogue. Our code and data is available at https://github.com/ankursikarwar/Cosmic.

空间理解多模态对话系统大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。