发现3D大模型可能依赖文字线索而非真正理解空间关系。
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
- 用纯文本微调模型在SQA3D上表现接近甚至超越3D-LLM。
- 新基准Real-3DQA显示现有模型在去除简单线索后性能大幅下降。
- 提出重加权训练目标,显著提升模型对3D视觉线索的依赖能力。
近期3D大语言模型声称能理解3D世界中的物体空间关系,但我们发现仅在纯文本问答对上微调语言模型,即可在SQA3D基准上达到相当或更优表现,且无需使用任何3D输入。这表明SQA3D基准可能无法识别模型是否依赖文本捷径而非进行3D感知推理。为解决此问题,我们提出更严格的评估基准Real-3DQA,通过过滤易猜题并引入结构化分类体系,评估多种3D推理维度。在Real-3DQA上的实验确认,一旦去除简单线索,现有3D-LLM在空间关系理解上表现不佳。我们进一步提出3D重加权训练目标,引导模型更多依赖3D视觉线索,显著提升其在空间推理任务中的性能。研究强调需建立稳健基准与定制训练策略,以推动真正的3D视觉-语言理解发展。
原文摘要 · Abstract (English)
Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on text-only question-answer pairs can perform comparably or even surpass these methods on the SQA3D benchmark without using any 3D input. This indicates that the SQA3D benchmark may not be able to detect if the model exploits textual shortcuts rather than engages in 3D-aware reasoning. To address this issue, we introduce Real-3DQA, a more rigorous evaluation benchmark that filters out easy-to-guess questions and introduces a structured taxonomy to assess various aspects of 3D reasoning. Experiments on Real-3DQA confirm that existing 3D-LLMs struggle with spatial relationships once simple cues are removed. We further propose a 3D-reweighted training objective that guides model to rely more on 3D visual clues, substantially enhancing 3D-LLMs performance in spatial reasoning tasks. Our findings underscore the need for robust benchmarks and tailored training strategies to advance genuine 3D vision-language understanding. Project page: https://real-3dqa.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。