3D视觉语言模型依赖语言线索多,3D编码器实际用得少。
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
- 对比3种3D模型架构,发现场景中心型模型效果最差
- 3D编码器在预训练阶段作用弱,大模型也难提升性能
- 新设计的问答数据集能防止模型走捷径,促进真理解
2D视觉语言模型(VLMs)的显著进展推动了其向3D场景的应用,用于3D问答、密集描述和视觉定位等任务。基于编码器设计,现有3D VLMs可分为以3D物体为中心、基于2D图像和以3D场景为中心三类。尽管场景中心型模型在架构上与2D模型相似,但性能远逊于最新的物体中心和图像基模型。深入分析发现,这类模型对3D场景编码器依赖度低,且预训练阶段有效性不如2D模型。此外,数据量增大带来的收益也不明显。研究显示,这些模型虽具备跨模态对齐能力,却过度依赖语言线索并过拟合常见答案分布,导致3D编码器未被有效利用。为此,我们提出一种新型3D相关性判别问答数据集,旨在打破捷径学习,提升3D理解能力。研究强调需改进评估方法与训练策略,以实现真正有效的3D视觉理解。
原文摘要 · Abstract (English)
Remarkable progress in 2D Vision-Language Models (VLMs) has spurred interest in extending them to 3D settings for tasks like 3D Question Answering, Dense Captioning, and Visual Grounding. Unlike 2D VLMs that typically process images through an image encoder, 3D scenes, with their intricate spatial structures, allow for diverse model architectures. Based on their encoder design, this paper categorizes recent 3D VLMs into 3D object-centric, 2D image-based, and 3D scene-centric approaches. Despite the architectural similarity of 3D scene-centric VLMs to their 2D counterparts, they have exhibited comparatively lower performance compared with the latest 3D object-centric and 2D image-based approaches. To understand this gap, we conduct an in-depth analysis, revealing that 3D scene-centric VLMs show limited reliance on the 3D scene encoder, and the pre-train stage appears less effective than in 2D VLMs. Furthermore, we observe that data scaling benefits are less pronounced on larger datasets. Our investigation suggests that while these models possess cross-modal alignment capabilities, they tend to over-rely on linguistic cues and overfit to frequent answer distributions, thereby diminishing the effective utilization of the 3D encoder. To address these limitations and encourage genuine 3D scene understanding, we introduce a novel 3D Relevance Discrimination QA dataset designed to disrupt shortcut learning and improve 3D understanding. Our findings highlight the need for advanced evaluation and improved strategies for better 3D understanding in 3D VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。