arXiv:2411.18728cs.CVcs.LG2024-11

仅用50个目标域标签,逼近监督学习效果

The Last Mile to Supervised Performance: Semi-Supervised Domain Adaptation for Semantic Segmentation

  • 融合一致性正则、像素对比学习与自训练
  • GTA到Cityscapes上达近监督性能
  • 适合低标注成本的语义分割场景

监督深度学习需要大量标注数据,但密集任务如语义分割的标注成本高且难获取。为此,研究者探索了无监督域适应(UDA)和半监督学习(SSL),但以低标注成本达到监督性能仍是难题。本文研究半监督域适应(SSDA)设置,提出一个简单框架,结合一致性正则、像素对比学习与自训练,有效利用少量目标域标签。在GTA-to-Cityscapes基准上优于现有方法,仅需50个目标标签即可实现近监督性能。在Synthia-to-Cityscapes、GTA-to-BDD和Synthia-to-BDD上的结果进一步验证了方法的有效性与实用性。此外,发现现有UDA与SSL方法不适用于SSDA设置,提出适配设计模式。

原文摘要 · Abstract (English)

Supervised deep learning requires massive labeled datasets, but obtaining annotations is not always easy or possible, especially for dense tasks like semantic segmentation. To overcome this issue, numerous works explore Unsupervised Domain Adaptation (UDA), which uses a labeled dataset from another domain (source), or Semi-Supervised Learning (SSL), which trains on a partially labeled set. Despite the success of UDA and SSL, reaching supervised performance at a low annotation cost remains a notoriously elusive goal. To address this, we study the promising setting of Semi-Supervised Domain Adaptation (SSDA). We propose a simple SSDA framework that combines consistency regularization, pixel contrastive learning, and self-training to effectively utilize a few target-domain labels. Our method outperforms prior art in the popular GTA-to-Cityscapes benchmark and shows that as little as 50 target labels can suffice to achieve near-supervised performance. Additional results on Synthia-to-Cityscapes, GTA-to-BDD and Synthia-to-BDD further demonstrate the effectiveness and practical utility of the method. Lastly, we find that existing UDA and SSL methods are not well-suited for the SSDA setting and discuss design patterns to adapt them.

语义分割半监督域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。