arXiv:2603.09774cs.AI2026-03

让大模型像人一样理解空间关系,无需训练即可推理三维场景。

World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models

  • 用3D重建和分割构建空间认知地图,生成结构化空间信息。
  • 在多个模型上提升5%~18%的空间推理准确率,纯文本模型也接近多模态效果。
  • 适合做空间推理、导航或智能体任务的研究者与开发者使用。

当前多模态基础模型在稳健的空间推理方面仍面临挑战。现有方法或因依赖3D标注数据而过拟合统计捷径,或局限于2D视觉感知,限制了在未见场景中的推理精度与泛化能力。受生物智能空间认知机制启发,我们提出World2Mind——一种无需训练的空间智能工具包。其核心利用3D重建与实例分割模型构建结构化空间认知地图,使基础模型能主动获取目标地标与路径的空间知识。为提供鲁棒的几何-拓扑先验,World2Mind合成一种以椭圆参数建模的异源空间树(Allocentric-Spatial Tree, AST),精准刻画地标布局。为缓解3D重建固有误差,引入三阶段推理链:工具调用评估、模态解耦线索收集、几何与语义交织推理。大量实验表明,World2Mind可使GPT-5.2等前沿模型性能提升5%~18%。令人惊讶的是,仅凭AST结构化文本,纯文本基础模型即可完成复杂3D空间推理,表现逼近先进多模态模型。

原文摘要 · Abstract (English)

Achieving robust spatial reasoning remains a fundamental challenge for current Multimodal Foundation Models (MFMs). Existing methods either overfit statistical shortcuts via 3D grounding data or remain confined to 2D visual perception, limiting both spatial reasoning accuracy and generalization in unseen scenarios. Inspired by the spatial cognitive mapping mechanisms of biological intelligence, we propose World2Mind, a training-free spatial intelligence toolkit. At its core, World2Mind leverages 3D reconstruction and instance segmentation models to construct structured spatial cognitive maps, empowering MFMs to proactively acquire targeted spatial knowledge regarding interested landmarks and routes of interest. To provide robust geometric-topological priors, World2Mind synthesizes an Allocentric-Spatial Tree (AST) that uses elliptical parameters to model the top-down layout of landmarks accurately. To mitigate the inherent inaccuracies of 3D reconstruction, we introduce a three-stage reasoning chain comprising tool invocation assessment, modality-decoupled cue collection, and geometry-semantics interwoven reasoning. Extensive experiments demonstrate that World2Mind boosts the performance of frontier models, such as GPT-5.2, by 5%~18%. Astonishingly, relying solely on the AST-structured text, purely text-only foundation models can perform complex 3D spatial reasoning, achieving performance approaching that of advanced multimodal models.

空间推理认知建模多模态3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。