arXiv:2504.14231cs.CV2025-04CVPR被引 3

利用图像点云联合信息,提升3D语义分割无监督域适应性能

Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation

  • 基于图像与点云双模态引导融合网络,动态调整特征权重
  • 在多个数据集上实现平均6.5 mIoU的性能提升
  • 适合需要跨域迁移的3D视觉任务研究者参考

视觉基础模型(VFMs)已成为图像分类、分割和目标定位等下游视觉任务的主流选择。本文进一步探索其在跨域3D语义分割中的应用,即从有标签源域向无标签目标域进行适应。方法利用配对的2D-3D数据(图像与点云),借助VFM提供的鲁棒跨域特征,在混合的有标签源数据与无标签目标数据上训练3D主干网络。核心是一个双模态引导的融合网络,根据目标域特性动态调节图像与点云流的贡献度。在多种设置下与当前最优方法对比,均取得显著性能提升,例如在所有任务上平均达到6.5 mIoU的增益。

原文摘要 · Abstract (English)

Vision Foundation Models (VFMs) have become a de facto choice for many downstream vision tasks, like image classification, image segmentation, and object localization. However, they can also provide significant utility for downstream 3D tasks that can leverage the cross-modal information (e.g., from paired image data). In our work, we further explore the utility of VFMs for adapting from a labeled source to unlabeled target data for the task of LiDAR-based 3D semantic segmentation. Our method consumes paired 2D-3D (image and point cloud) data and relies on the robust (cross-domain) features from a VFM to train a 3D backbone on a mix of labeled source and unlabeled target data. At the heart of our method lies a fusion network that is guided by both the image and point cloud streams, with their relative contributions adjusted based on the target domain. We extensively compare our proposed methodology with different state-of-the-art methods in several settings and achieve strong performance gains. For example, achieving an average improvement of 6.5 mIoU (over all tasks), when compared with the previous state-of-the-art.

3D分割域适应多模态融合视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。