arXiv:2504.04540cs.CVcs.AI2025-04

对比点云、视觉与文本,发现点云能提升空间推理能力但仍有局限。

The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study

  • 构建多模态3D推理基准ScanReQA,统一评估不同模态表现
  • 点云与视觉模型在空间推理上优于纯文本模型,但二元推理仍困难
  • 发现3D LLM存在注意力坍缩现象,影响空间理解能力

基于点云的3D大语言模型在空间推理任务中备受关注,但其相比其他模态的优势尚不明确。现有3D基准也难以公平评估多模态大模型对空间概念的理解能力。为此,我们提出ScanReQA,一个融合文本、视觉与点云三模态的3D空间推理基准。通过在该基准上评估文本、2D与3D LLMs的表现,我们比较了不同模态在理解空间概念上的有效性。进一步分析3D LLMs的推理机制后发现:1)当前3D LLMs在二元空间推理任务中仍具挑战性;2)基于点云与视觉的多模态模型比纯文本模型具备更强的空间推理能力;3)3D LLMs表现出类似2D LLMs的注意力坍缩现象,削弱了空间推理性能。这些结论为后续3D LLMs发展及跨模态基础模型研究提供重要启示。数据集与代码已开源:https://github.com/EmbodiedCity/ScanReQA.code。

原文摘要 · Abstract (English)

3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some promising results, the advantages of point clouds over other modalities remain unclear. Moreover, existing 3D benchmarks are insufficient for fairly evaluating the ability of multimodal LLMs to comprehend spatial concepts. To address these challenges, we introduce ScanReQA, a 3D spatial reasoning benchmark encompassing text, vision, and point cloud modalities. We then evaluate the performance of text, 2D, and 3D LLMs on the benchmark to compare the effectiveness of different modalities in understanding spatial concepts. Furthermore, we analyze the reasoning mechanisms behind 3D LLMs using point clouds. Our findings reveal that: 1) binary spatial reasoning remains challenging for current 3D LLMs, 2) MLLMs based on point cloud and visual modalities demonstrate stronger spatial reasoning capabilities than LLMs, and 3) 3D LLMs exhibit the attention sink phenomenon similar to that in 2D LLMs, impairing spatial reasoning. We think these conclusions can help the next step of 3D LLMs and also offer insights for foundation models in other modalities. We release datasets and codes in the project page: https://github.com/EmbodiedCity/ScanReQA.code.

3D LLM空间推理点云多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。