通过去噪提升多视角融合,让视觉语言模型更好理解开放词汇的3D场景。
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding
- 利用区域级图像与文本特征,结合3D几何先验优化多视角特征融合
- 在ScanNet200和Matterport160上分别达到14.7%和16.2%的mIoU新纪录
- 无需训练即可降低视觉语言模型噪声,适合开放词汇3D理解任务
当前开放词汇3D场景理解方法主要依赖对比学习或2D特征蒸馏,但在多样化物体类别上表现受限,因3D数据量不足难以训练强泛化能力的模型。我们发现2D多视角融合在理解复杂3D概念上更具优势,但受视觉语言模型固有噪声影响,性能未达最优。为此,提出MVOV3D,旨在释放2D多视角融合潜力。该方法不通过训练消除噪声,而是利用精确的区域级图像与文本特征(来自CLIP编码器)并结合3D几何先验,优化多视角融合过程。大量实验表明,MVOV3D在多个数据集上表现优异:在ScanNet200上实现14.7% mIoU,Matterport160上达到16.2% mIoU,显著超越现有主流3D网络,验证了其在开放词汇语义分割中的有效性。
原文摘要 · Abstract (English)
Recent open-vocabulary 3D scene understanding approaches mainly focus on training 3D networks through contrastive learning with point-text pairs or by distilling 2D features into 3D models via point-pixel alignment. While these methods show considerable performance in benchmarks with limited vocabularies, they struggle to handle diverse object categories as the limited amount of 3D data upbound training strong open-vocabulary 3d models. We observe that 2D multi-view fusion methods take precedence in understanding diverse concepts in 3D scenes. However, inherent noises in vision-language models lead multi-view fusion to sub-optimal performance. To this end, we introduce MVOV3D, a novel approach aimed at unleashing the potential of 2D multi-view fusion for open-vocabulary 3D scene understanding. We focus on reducing the inherent noises without training, thereby preserving the generalizability while enhancing open-world capabilities. Specifically, MVOV3D improves multi-view 2D features by leveraging precise region-level image features and text features encoded by CLIP encoders and incorporates 3D geometric priors to optimize multi-view fusion. Extensive experiments on various datasets demonstrate the effectiveness of our method. Notably, our MVOV3D achieves a new record with 14.7% mIoU on ScanNet200 and 16.2% mIoU on Matterport160 for challenge open-vocabulary semantic segmentation, outperforming current leading trained 3D networks by a significant margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。