研究发现,用真实数据训练的视觉大模型无需无监督域适应也能表现很好。
What is the Added Value of UDA in the VFM Era?
- 对比源域仅微调与无监督域适应在不同数据场景下的效果
- 使用强合成数据时,域适应提升从+8 mIoU降至+2 mIoU
- 少量标注目标数据下,合成域适应可达到全监督水平(85 mIoU)
无监督域适应(UDA)能提升感知模型在未标注目标域上的泛化能力。当使用视觉基础模型(VFMs)和合成源数据时,其性能可媲美完全监督学习。然而,由于VFMs具备强大预训练泛化能力,仅对源域进行微调也表现良好。当前学术研究中的数据场景未必反映真实应用,因此尚不明确:(a) UDA在更真实、多样化的数据下表现如何;(b) VFMs的源域微调是否同样有效。本研究以语义分割为例,评估了合成到真实和真实到真实两种场景下的UDA效果,还考察了少量标注目标数据的影响。结果表明,尽管场景更真实,但并不一定更难。当使用更强合成数据时,UDA相对于源域微调的提升从+8 mIoU降至+2 mIoU;当使用更多样真实源数据时,UDA无额外价值。但在所有合成数据场景中,UDA始终优于源域微调;即使仅使用1/16的Cityscapes标签,合成UDA仍能达到85 mIoU的顶尖性能,与全监督模型相当。综合结果,我们探讨了如何在大规模自动驾驶中有效利用UDA。
原文摘要 · Abstract (English)
Unsupervised Domain Adaptation (UDA) can improve a perception model's generalization to an unlabeled target domain starting from a labeled source domain. UDA using Vision Foundation Models (VFMs) with synthetic source data can achieve generalization performance comparable to fully-supervised learning with real target data. However, because VFMs have strong generalization from their pre-training, more straightforward, source-only fine-tuning can also perform well on the target. As data scenarios used in academic research are not necessarily representative for real-world applications, it is currently unclear (a) how UDA behaves with more representative and diverse data and (b) if source-only fine-tuning of VFMs can perform equally well in these scenarios. Our research aims to close these gaps and, similar to previous studies, we focus on semantic segmentation as a representative perception task. We assess UDA for synth-to-real and real-to-real use cases with different source and target data combinations. We also investigate the effect of using a small amount of labeled target data in UDA. We clarify that while these scenarios are more realistic, they are not necessarily more challenging. Our results show that, when using stronger synthetic source data, UDA's improvement over source-only fine-tuning of VFMs reduces from +8 mIoU to +2 mIoU, and when using more diverse real source data, UDA has no added value. However, UDA generalization is always higher in all synthetic data scenarios than source-only fine-tuning and, when including only 1/16 of Cityscapes labels, synthetic UDA obtains the same state-of-the-art segmentation quality of 85 mIoU as a fully-supervised model using all labels. Considering the mixed results, we discuss how UDA can best support robust autonomous driving at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。