融合多模型语义与几何能力,提升3D场景开放词汇理解性能
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
- 集成CLIP、DINOv2等多模型特征,统一跨模态表示
- 在ScanNetV2和Matterport3D上实现领先分割效果
- 适合做3D视觉与多模态融合的研究者参考
缺乏大规模3D-文本语料库导致近期工作依赖视觉语言模型(VLMs)蒸馏开放词汇知识。但这些方法通常仅用单一VLM对齐3D模型的特征空间,限制了3D模型利用不同基础模型所蕴含的多样空间与语义能力。本文提出首个整合多个基础模型(如CLIP、DINOv2、Stable Diffusion)的开放词汇3D场景理解框架CUA-O3D。引入确定性不确定性估计,自适应地蒸馏并调和来自各模型的异构2D特征嵌入。解决两大挑战:(1)融合VLM的语义先验与空间感知模型的几何知识;(2)通过新颖的确定性不确定性估计捕捉不同模型在语义与几何敏感度上的特异性,促进训练中异构表示的融合。在ScanNetV2和Matterport3D上的实验表明,该方法不仅显著提升开放词汇分割性能,还实现稳健的跨域对齐与竞争力的空间感知能力。
原文摘要 · Abstract (English)
The lack of a large-scale 3D-text corpus has led recent works to distill open-vocabulary knowledge from vision-language models (VLMs). However, these methods typically rely on a single VLM to align the feature spaces of 3D models within a common language space, which limits the potential of 3D models to leverage the diverse spatial and semantic capabilities encapsulated in various foundation models. In this paper, we propose Cross-modal and Uncertainty-aware Agglomeration for Open-vocabulary 3D Scene Understanding dubbed CUA-O3D, the first model to integrate multiple foundation models-such as CLIP, DINOv2, and Stable Diffusion-into 3D scene understanding. We further introduce a deterministic uncertainty estimation to adaptively distill and harmonize the heterogeneous 2D feature embeddings from these models. Our method addresses two key challenges: (1) incorporating semantic priors from VLMs alongside the geometric knowledge of spatially-aware vision foundation models, and (2) using a novel deterministic uncertainty estimation to capture model-specific uncertainties across diverse semantic and geometric sensitivities, helping to reconcile heterogeneous representations during training. Extensive experiments on ScanNetV2 and Matterport3D demonstrate that our method not only advances open-vocabulary segmentation but also achieves robust cross-domain alignment and competitive spatial perception capabilities. The code will be available at: https://github.com/TyroneLi/CUA_O3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。