arXiv:2601.11442cs.CVcs.AI2026-01被引 5

用度量认知地图实现3D视觉模型的可解释空间推理

Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps

  • 构建度量认知地图融合离散关系与连续几何表示
  • 仅用一半标注数据达59.9%准确率,接近全量训练60.9%水平
  • 在小样本场景下显著优于当前最优方法,适合可解释性需求强的场景

我们提出Map2Thought框架,为3D视觉语言模型提供显式且可解释的空间推理能力。该框架基于两个核心组件:度量认知地图(Metric-CogMap)和认知思维链(Cog-CoT)。Metric-CogMap通过融合离散网格的关系推理能力与连续度量尺度的几何理解能力,构建统一的空间表征。在此基础上,Cog-CoT通过确定性操作(如向量运算、边界框距离计算、遮挡感知的外观顺序线索)进行显式几何推理,生成基于3D结构的可解释推理轨迹。实验表明,Map2Thought实现了可解释的3D理解,在仅使用一半监督信号的情况下达到59.9%的准确率,接近全数据集训练的60.9%基准表现。在VSI-Bench数据集上,于10%、25%、50%训练子集下,分别领先当前最优方法5.3%、4.8%、4.0%,展现出卓越的小样本泛化能力。

原文摘要 · Abstract (English)

We propose Map2Thought, a framework that enables explicit and interpretable spatial reasoning for 3D VLMs. The framework is grounded in two key components: Metric Cognitive Map (Metric-CogMap) and Cognitive Chain-of-Thought (Cog-CoT). Metric-CogMap provides a unified spatial representation by integrating a discrete grid for relational reasoning with a continuous, metric-scale representation for precise geometric understanding. Building upon the Metric-CogMap, Cog-CoT performs explicit geometric reasoning through deterministic operations, including vector operations, bounding-box distances, and occlusion-aware appearance order cues, producing interpretable inference traces grounded in 3D structure. Experimental results show that Map2Thought enables explainable 3D understanding, achieving 59.9% accuracy using only half the supervision, closely matching the 60.9% baseline trained with the full dataset. It consistently outperforms state-of-the-art methods by 5.3%, 4.8%, and 4.0% under 10%, 25%, and 50% training subsets, respectively, on the VSI-Bench.

3D理解可解释性空间推理小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。