arXiv:2511.01618cs.CVcs.CL2025-11NeurIPS被引 9

让多模态大模型学会看懂物体的3D空间关系,提升跨视角理解能力。

Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models

论文配图:Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
图 1 · 摘自论文原文
  • 设计新任务与数据集,训练模型理解不同视角下的物体空间关系。
  • 在10万组图像对上微调后,跨域推理任务准确率显著提升。
  • 适合研究机器人视觉、三维场景理解及多模态模型泛化性的团队。

多模态大语言模型(MLLMs)在2D视觉理解方面取得显著进展,激发了其在复杂3D推理任务中的应用兴趣。然而,这些模型是否能有效捕捉真实世界中所需的详细空间信息,特别是跨视角一致性——3D推理的关键要求——仍不明确。为此,我们提出视点学习(Viewpoint Learning)任务,用于评估和提升MLLM的空间推理能力。我们构建了包含10万组以物体为中心的图像对及其对应问答对的Viewpoint-100K数据集。方法采用两阶段微调:首先在该数据集上通过监督微调(SFT)注入基础空间知识,显著提升多个任务表现;其次在更广泛问题集上使用分组相对策略优化(GRPO)进行强化学习,增强泛化能力。此外,我们提出一种混合冷启动初始化方法,可同时学习视点表征并保持连贯推理。实验表明,该方法显著激活了MLLM的空间推理能力,在域内与域外推理任务中均有明显提升。结果强调为MLLM发展基础空间技能的重要性,有助于推动机器人、自动驾驶系统及3D场景理解的发展。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can effectively capture the detailed spatial information required for robust real-world performance, especially cross-view consistency, a key requirement for accurate 3D reasoning. Considering this issue, we introduce Viewpoint Learning, a task designed to evaluate and improve the spatial reasoning capabilities of MLLMs. We present the Viewpoint-100K dataset, consisting of 100K object-centric image pairs with diverse viewpoints and corresponding question-answer pairs. Our approach employs a two-stage fine-tuning strategy: first, foundational knowledge is injected to the baseline MLLM via Supervised Fine-Tuning (SFT) on Viewpoint-100K, resulting in significant improvements across multiple tasks; second, generalization is enhanced through Reinforcement Learning using the Group Relative Policy Optimization (GRPO) algorithm on a broader set of questions. Additionally, we introduce a hybrid cold-start initialization method designed to simultaneously learn viewpoint representations and maintain coherent reasoning thinking. Experimental results show that our approach significantly activates the spatial reasoning ability of MLLM, improving performance on both in-domain and out-of-domain reasoning tasks. Our findings highlight the value of developing foundational spatial skills in MLLMs, supporting future progress in robotics, autonomous systems, and 3D scene understanding.

多模态空间推理3D理解视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。