arXiv:2412.16859cs.CVcs.AI2024-12CVPR被引 3

用生成模型对齐虚拟与真实图像特征,提升无监督分割效果

Adversarially Domain-adaptive Latent Diffusion for Unsupervised Semantic Segmentation

  • 通过扩散模型构建跨域特征连接,增强上下文理解
  • 在GTA5→Cityscapes上达74.4 mIoU,优于现有方法
  • 适合做合成数据到真实场景的语义分割迁移

语义分割依赖大量像素级标注,无监督域适应(UDA)旨在将标注源域知识迁移到未标注或弱标注目标域。利用可控虚拟环境(如视频游戏或交通模拟器)生成的合成数据可自动提供像素级标注。然而,即使存在此类数据,学习能同时捕捉虚拟与真实世界差异的泛化表征仍具挑战,源于二者在概率分布和几何结构上的不一致。本文提出基于潜在扩散模型的语义分割方法——跨编码器连接潜在扩散模型(ICCLD),结合无监督域适应策略。模型通过跨编码器连接提升上下文理解并保留细节,同时利用对抗学习在潜在扩散过程中对齐跨域特征分布。在GTA5、Synthia和Cityscapes上的实验表明,ICCLD超越当前最优的UDA方法,在GTA5→Cityscapes任务中达到74.4 mIoU,Synthia→Cityscapes任务中达67.2 mIoU。

原文摘要 · Abstract (English)

Semantic segmentation requires extensive pixel-level annotation, motivating unsupervised domain adaptation (UDA) to transfer knowledge from labelled source domains to unlabelled or weakly labelled target domains. One of the most efficient strategies involves using synthetic datasets generated within controlled virtual environments, such as video games or traffic simulators, which can automatically generate pixel-level annotations. However, even when such datasets are available, learning a well-generalised representation that captures both domains remains challenging, owing to probabilistic and geometric discrepancies between the virtual world and real-world imagery. This work introduces a semantic segmentation method based on latent diffusion models, termed Inter-Coder Connected Latent Diffusion (ICCLD), alongside an unsupervised domain adaptation approach. The model employs an inter-coder connection to enhance contextual understanding and preserve fine details, while adversarial learning aligns latent feature distributions across domains during the latent diffusion process. Experiments on GTA5, Synthia, and Cityscapes demonstrate that ICCLD outperforms state-of-the-art UDA methods, achieving mIoU scores of 74.4 (GTA5$\rightarrow$Cityscapes) and 67.2 (Synthia$\rightarrow$Cityscapes).

语义分割域自适应扩散模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。