arXiv:2503.10685cs.CVeess.IV2025-03ICCV被引 2

用视觉大模型提升无监督域适应分割性能,优化多尺度特征与损失函数。

VFM-UDA++: Improving Network Architectures and Data Strategies for Unsupervised Domain Adaptive Semantic Segmentation

  • 引入多尺度特征与适配ViT的损失函数,改进模型结构。
  • 在GTA5→Cityscapes上提升1.4 mIoU,数据量增加时再增2.4 mIoU。
  • 适合关注大模型与数据扩展协同效应的研究者。

无监督域适应(UDA)可使模型从有标签源域泛化到无标签目标域,尤其在数据有限时表现突出。同时,大规模无监督预训练的视觉基础模型(VFMs)也展现出优异的下游性能和泛化能力。这促使我们探索如何让UDA更好利用VFMs。已有工作(VFM-UDA)表明,用VFM替换ImageNet预训练编码器能提升泛化能力,但发现常用特征距离损失在应用于基于ViT的VFM时会损害性能。此外,该方法未引入多尺度归纳偏置,而后者已被证明对语义分割有益。为此,本文提出VFM-UDA++:(1)研究多尺度特征的作用;(2)设计适配ViT-VFM的特征距离损失;(3)评估合成源数据与真实目标数据增加对性能的影响。通过上述改进,在标准的GTA5→Cityscapes基准上实现+1.4 mIoU提升。不同于传统非VFM UDA方法无法随数据增长而提升,VFM-UDA++表现出持续增益,当数据规模扩大时额外获得+2.4 mIoU,表明基于VFM的UDA仍能从更多数据中受益。

原文摘要 · Abstract (English)

Unsupervised Domain Adaptation (UDA) enables strong generalization from a labeled source domain to an unlabeled target domain, often with limited data. In parallel, Vision Foundation Models (VFMs) pretrained at scale without labels have also shown impressive downstream performance and generalization. This motivates us to explore how UDA can best leverage VFMs. Prior work (VFM-UDA) demonstrated that replacing a standard ImageNet-pretrained encoder with a VFM improves generalization. However, it also showed that commonly used feature distance losses harm performance when applied to VFMs. Additionally, VFM-UDA does not incorporate multi-scale inductive biases, which are known to improve semantic segmentation. Building on these insights, we propose VFM-UDA++, which (1) investigates the role of multi-scale features, (2) adapts feature distance loss to be compatible with ViT-based VFMs and (3) evaluates how UDA benefits from increased synthetic source and real target data. By addressing these questions, we can improve performance on the standard GTA5 $\rightarrow$ Cityscapes benchmark by +1.4 mIoU. While prior non-VFM UDA methods did not scale with more data, VFM-UDA++ shows consistent improvement and achieves a further +2.4 mIoU gain when scaling the data, demonstrating that VFM-based UDA continues to benefit from increased data availability.

无监督域适应视觉大模型语义分割多尺度特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。