无需标注数据,让多模态感知自动融合视觉与深度信息。
UMSS: Towards Unsupervised Multi-modal Semantic Segmentation

- 构建统一潜在空间,通过跨模态对应协同学习共享语义
- 在NYU Depth v2和MFNet上分别提升mIoU 6.4%和9.8%
- 适合无监督多模态分割任务的研究者与工业应用
多模态语义分割对复杂环境下的鲁棒感知至关重要,但其发展受限于人工标注的高昂成本。尽管单模态无监督分割已取得良好效果,但直接扩展至多模态数据常因融合退化而受阻,原因在于缺乏显式监督时,不同传感器捕捉的异构结构模式难以协调,无法有效利用互补信息。本文首次提出无监督多模态语义分割(UMSS)问题,并提出UniM2框架,基于DINOv3将传统融合方法转化为稳定性能提升。核心思想是通过跨模态对应协同(CMCS)驱动统一潜在空间,提取内在共享语义线索,避免依赖标签引导的自适应融合。为缓解模态间冲突,引入跨模态调谐器(CMH),以RGB为稳定参考,抑制不一致关系监督,引导模型挖掘互补结构特征。在NYU Depth v2和MFNet上的实验表明,UniM2分别实现6.4%和9.8%的mIoU提升,显著优于现有框架。
原文摘要 · Abstract (English)
Multimodal semantic segmentation (MSS) is essential for robust perception in complex environments, yet its potential remains largely untapped because of the prohibitive cost of human annotations. While unsupervised semantic segmentation (USS) has achieved strong results on a single RGB modality, its naive extension to multimodal data is often hindered by fusion degradation. This occurs because, without explicit supervision, existing frameworks struggle to reconcile the heterogeneous structural patterns captured by different sensors and therefore fail to effectively exploit their complementary information. In this paper, we make the first attempt to address the novel problem of Unsupervised Multimodal Semantic Segmentation (UMSS), aiming to effectively exploit complementary sensor information in a fully label free setting. To this end, we propose UniM2 (Unified Multimodal), a novel framework built on DINOv3 that transforms conventional fusion methods into consistent performance gains. Our key idea is to learn a unified latent space driven by Cross Modal Correspondence Synergy (CMCS) to extract intrinsic shared semantic cues, bypassing the need for label guided adaptive fusion. To mitigate inherent intermodal conflicts, we introduce a Cross Modal Harmonizer (CMH) that designates RGB as a stable reference, effectively suppressing inconsistent relational supervision while guiding the model to exploit complementary structural features. Extensive experimental results on NYU Depth v2 and MFNet show that UniM2 improves mIoU by 6.4% and 9.8%, respectively, demonstrating clear advantages over existing frameworks for UMSS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。