arXiv:2504.20024cs.CV2025-04NeurIPS被引 66

让AI显式理解3D空间关系,提升推理能力与泛化性

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning

论文配图:SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
图 1 · 摘自论文原文
  • 引入显式3D表征贯穿感知、计算和推理全流程
  • 在3DSRBench上超越Gemini 2.0 9.2%,新问题泛化更强
  • 适合研究3D视觉推理与模型可解释性的开发者

尽管多模态模型取得进展,当前开源与专有模型在3D空间推理上仍面临挑战。现有方法通过微调3D视觉问答数据提升性能,但多采用隐式推理,面对人类轻易解答的问题仍失败,即使使用长链式思维。本文提出SpatialReasoner,一种新型大视觉语言模型(LVLM),通过在多个阶段共享显式3D表示,实现3D空间推理的显式建模。该设计提供连贯接口,支持高级3D空间推理并增强对新类型问题的泛化能力。通过对多步推理轨迹中的显式3D表示分析,我们揭示了当前LVLM的事实错误及关键缺陷。结果表明,SpatialReasoner在多种空间推理基准上表现更优,在3DSRBench上优于Gemini 2.0 9.2%,且在新问题评估中展现更强泛化能力。本研究将先前视觉基础模型的3D解析能力与大语言模型的强大推理能力结合,开辟3D空间推理新方向。

原文摘要 · Abstract (English)

Despite recent advances on multi-modal models, 3D spatial reasoning remains a challenging task for state-of-the-art open-source and proprietary models. Recent studies explore data-driven approaches and achieve enhanced spatial reasoning performance by fine-tuning models on 3D-related visual question-answering data. However, these methods typically perform spatial reasoning in an implicit manner and often fail on questions that are trivial to humans, even with long chain-of-thought reasoning. In this work, we introduce SpatialReasoner, a novel large vision-language model (LVLM) that addresses 3D spatial reasoning with explicit 3D representations shared between multiple stages--3D perception, computation, and reasoning. Explicit 3D representations provide a coherent interface that supports advanced 3D spatial reasoning and improves the generalization ability to novel question types. Furthermore, by analyzing the explicit 3D representations in multi-step reasoning traces of SpatialReasoner, we study the factual errors and identify key shortcomings of current LVLMs. Results show that our SpatialReasoner achieves improved performance on a variety of spatial reasoning benchmarks, outperforming Gemini 2.0 by 9.2% on 3DSRBench, and generalizes better when evaluating on novel 3D spatial reasoning questions. Our study bridges the 3D parsing capabilities of prior visual foundation models with the powerful reasoning abilities of large language models, opening new directions for 3D spatial reasoning.

3D推理视觉语言模型显式表征泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。