arXiv:2603.10370cs.CV2026-03

让多模态模型学会自主判断何时需要几何信息,提升空间推理能力。

GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning

  • 模型自动生成几何感知判断,仅在2D信息不足时调用几何特征。
  • 在多个空间推理基准上性能显著提升,且不损失原有视觉理解能力。
  • 适合需要高效精准空间推理的智能系统研发人员。

迈向人工超智能需要强大的感知能力。当前多模态大模型在空间理解上存在局限,几何信息至关重要。现有方法通常强制注入几何信号,忽视其必要性并增加计算开销。本文提出一种新框架,使模型具备感知不足的自知能力,仅当二维线索不足时才自主启用几何特征进行推理。首先,在模型架构中引入独立几何输入通道并进行对齐训练,有效利用几何信息;其次,构建专用的空间感知监督微调数据集,激活模型内部隐含的感知提示,使其能自主判断几何信息的必要性。在多个空间推理基准上的实验验证了该方法的有效性,显著提升空间推理性能,同时保持原有2D视觉推理能力,为更鲁棒、高效、自我意识的多模态智能提供新路径。

原文摘要 · Abstract (English)

Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understanding of Multimodal Large Language Models (MLLMs), where geometry information is essential. Existing methods often address this by rigidly injecting geometric signals into every input, while ignoring their necessity and adding computation overhead. Contrary to this paradigm, our framework endows the model with an awareness of perceptual insufficiency, empowering it to autonomously engage geometric features in reasoning when 2D cues are deemed insufficient. To achieve this, we first introduce an independent geometry input channel to the model architecture and conduct alignment training, enabling the effective utilization of geometric features. Subsequently, to endow the model with perceptual awareness, we curate a dedicated spatial-aware supervised fine-tuning dataset. This serves to activate the model's latent internal cues, empowering it to autonomously determine the necessity of geometric information. Experiments across multiple spatial reasoning benchmarks validate this approach, demonstrating significant spatial gains without compromising 2D visual reasoning capabilities, offering a path toward more robust, efficient and self-aware multi-modal intelligence.

多模态空间推理几何感知自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。