arXiv:2506.23120cs.CV2025-06ICCV被引 7

通过分步推理提升多模态大模型的3D空间理解能力

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

  • 将空间推理拆分为找相关元素和结合视觉先验处理指令两步
  • 在2.5万训练样本上实现更精准的3D目标定位与理解
  • 适合需要复杂空间逻辑的机器人、自动驾驶场景

近年来,点云感知借助大语言模型(LLM)实现视觉-语言对齐,在场景理解方面取得显著进展。然而,面对需要精确空间推理的复杂指令时,现有方法仍存在挑战,即使3D点云数据已提供尺寸、位置等详细空间线索。为此,我们提出基于推理的分割框架R$^2$S,模拟人类认知过程:先识别相关要素,再结合其视觉先验处理指令。此外,鉴于现有数据集在复杂推理任务上的不足,我们构建了3D ReasonSeg数据集,包含25,185个训练样本和3,966个验证样本,具备精细标注。定量与定性实验表明,R$^2$S与3D ReasonSeg有效提升了点云感知的空间推理能力,可作为未来研究的新基线与基准。

原文摘要 · Abstract (English)

Recent advances in point cloud perception have demonstrated remarkable progress in scene understanding through vision-language alignment leveraging large language models (LLMs). However, existing methods may still encounter challenges in handling complex instructions that require accurate spatial reasoning, even if the 3D point cloud data provides detailed spatial cues such as size and position for identifying the targets. To tackle this issue, we propose Relevant Reasoning Segmentation (R$^2$S), a reasoning-based segmentation framework. The framework emulates human cognitive processes by decomposing spatial reasoning into two sequential stages: first identifying relevant elements, then processing instructions guided by their associated visual priors. Furthermore, acknowledging the inadequacy of existing datasets in complex reasoning tasks, we introduce 3D ReasonSeg, a reasoning-based segmentation dataset comprising 25,185 training samples and 3,966 validation samples with precise annotations. Both quantitative and qualitative experiments demonstrate that the R$^2$S and 3D ReasonSeg effectively endow 3D point cloud perception with stronger spatial reasoning capabilities, and we hope that they can serve as a new baseline and benchmark for future work.

空间推理点云理解多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。