提出分层掩码机制,提升语义分割无监督域适应性能
OMUDA: Omni-level Masking for Unsupervised Domain Adaptation in Semantic Segmentation
- 通过上下文、特征、类别三层次掩码策略,自适应区分前景背景
- 在SYNTHIA→Cityscapes等任务上平均提升7%,达当前最优
- 特别适合处理伪标签噪声和跨域特征不一致问题
无监督域适应(UDA)使语义分割模型能从有标签源域推广到无标签目标域。然而,现有方法仍受限于跨域上下文模糊、特征表示不一致及类别级伪标签噪声。为此,本文提出面向无监督域适应的全层次掩码框架(OMUDA),引入多层级掩码策略:1)上下文感知掩码(CAM)自适应区分前景与背景,平衡全局上下文与局部细节;2)特征蒸馏掩码(FDM)通过预训练模型知识迁移增强鲁棒一致的特征学习;3)类别解耦掩码(CDM)显式建模类别不确定性,缓解伪标签噪声影响。该分层掩码范式有效降低了上下文、表征与类别层面的域偏移,提供统一解决方案。在多个挑战性跨域语义分割基准上验证了有效性,尤其在SYNTHIA→Cityscapes与GTA5→Cityscapes任务中,可无缝集成至现有方法并持续取得领先效果,平均提升7%。
原文摘要 · Abstract (English)
Unsupervised domain adaptation (UDA) enables semantic segmentation models to generalize from a labeled source domain to an unlabeled target domain. However, existing UDA methods still struggle to bridge the domain gap due to cross-domain contextual ambiguity, inconsistent feature representations, and class-wise pseudo-label noise. To address these challenges, we propose Omni-level Masking for Unsupervised Domain Adaptation (OMUDA), a unified framework that introduces hierarchical masking strategies across distinct representation levels. Specifically, OMUDA comprises: 1) a Context-Aware Masking (CAM) strategy that adaptively distinguishes foreground from background to balance global context and local details; 2) a Feature Distillation Masking (FDM) strategy that enhances robust and consistent feature learning through knowledge transfer from pre-trained models; and 3) a Class Decoupling Masking (CDM) strategy that mitigates the impact of noisy pseudo-labels by explicitly modeling class-wise uncertainty. This hierarchical masking paradigm effectively reduces the domain shift at the contextual, representational, and categorical levels, providing a unified solution beyond existing approaches. Extensive experiments on multiple challenging cross-domain semantic segmentation benchmarks validate the effectiveness of OMUDA. Notably, on the SYNTHIA->Cityscapes and GTA5->Cityscapes tasks, OMUDA can be seamlessly integrated into existing UDA methods and consistently achieving state-of-the-art results with an average improvement of 7%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。