arXiv:2601.11779cs.CV2026-01被引 27

用无监督图像转换生成目标域数据,提升跨域目标检测性能

Cross-Domain Object Detection Using Unsupervised Image Translation

  • 仅用源域标注数据和目标域无标注数据,通过图像翻译生成目标域训练集
  • 在自动驾驶真实场景中显著超越现有方法,接近有目标域标签时的上限性能
  • 方法更简洁易懂,兼具高效性与可解释性,适合实际部署

无监督域适应中的目标检测旨在将源域训练的检测器适配到未见的目标域。近期方法通过对齐中间特征取得良好效果,但实现复杂且难以解释。本文提出一种新方法:利用两个无监督图像翻译模型(CycleGAN 和基于 AdaIN 的模型),仅使用源域标注数据和目标域无标注数据,生成目标域人工训练集并训练检测器。该方法在自动驾驶真实场景中表现优异,多数情况下超越当前最优方法,进一步缩小与使用目标域标注数据训练的上限性能差距。核心贡献在于提出了一种更简单、更有效且更具可解释性的跨域检测方案。

原文摘要 · Abstract (English)

Unsupervised domain adaptation for object detection addresses the adaption of detectors trained in a source domain to work accurately in an unseen target domain. Recently, methods approaching the alignment of the intermediate features proven to be promising, achieving state-of-the-art results. However, these methods are laborious to implement and hard to interpret. Although promising, there is still room for improvements to close the performance gap toward the upper-bound (when training with the target data). In this work, we propose a method to generate an artificial dataset in the target domain to train an object detector. We employed two unsupervised image translators (CycleGAN and an AdaIN-based model) using only annotated data from the source domain and non-annotated data from the target domain. Our key contributions are the proposal of a less complex yet more effective method that also has an improved interpretability. Results on real-world scenarios for autonomous driving show significant improvements, outperforming state-of-the-art methods in most cases, further closing the gap toward the upper-bound.

跨域检测图像翻译无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。