让AI通过3D建模看多视角,提升空间推理能力
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
- 用单图生成高精度3D网格,结合关键词提取和分层掩码
- 自动计算最优相机位姿,合成新视角,模拟人类换位思考
- 无需训练,在3DSRBench等数据集上超越GPT-5.2等模型
尽管多模态大语言模型取得了显著进展,但在复杂3D空间推理方面仍因依赖2D视觉先验而表现不足。现有方法要么依赖计算昂贵的有限3D数据微调,要么采用缺乏几何理解与视角灵活性的固定工具调用机制。为此,我们提出一种无需训练的框架,基于显式3D重建引入视觉思维链机制。该流程首先利用MLLM引导的关键词提取与多粒度掩码生成,从单张图像重建高保真3D网格;随后借助外部知识库迭代计算最优相机外参并合成新视角,模拟人类视角转换。大量实验表明,该方法显著增强空间理解能力:在3DSRBench和Rel3D等主流基准上,性能超越专用空间模型及通用MLLM(如GPT-5.2、Gemini-2.5-Flash)。
原文摘要 · Abstract (English)
Although Multimodal Large Language Models have achieved remarkable progress, they still struggle with complex 3D spatial reasoning due to the reliance on 2D visual priors. Existing approaches typically mitigate this limitation either through computationally expensive post-training procedures on limited 3D datasets or through rigid tool-calling mechanisms that lack explicit geometric understanding and viewpoint flexibility. To address these challenges, we propose a \textit{training-free} framework that introduces a Visual Chain-of-Thought mechanism grounded in explicit 3D reconstruction. The proposed pipeline first reconstructs a high-fidelity 3D mesh from a single image using MLLM-guided keyword extraction and mask generation at multiple granularities. Subsequently, the framework leverages an external knowledge base to iteratively compute optimal camera extrinsic parameters and synthesize novel views, thereby emulating human perspective-taking. Extensive experiments demonstrate that the proposed approach significantly enhances spatial comprehension. Specifically, the framework outperforms specialized spatial models and general-purpose MLLMs, including \textit{GPT-5.2} and \textit{Gemini-2.5-Flash}, on major benchmarks such as 3DSRBench and Rel3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。