arXiv:2506.17931cs.CVcs.AI2025-06中稿 · ICPR'24

提升自然图像域适应性能,解决多模态分布下的尺度与风格差异问题。

IDAL: Improved Domain Adaptive Learning for Natural Images Dataset

  • 融合ResNet与特征金字塔结构,分离处理内容与风格特征。
  • 设计组合损失函数,在多个数据集上实现更优准确率与更快收敛速度。
  • 适合处理具有复杂视觉变化的跨域图像识别任务,如办公、场景等场景迁移。

我们提出一种针对自然图像的无监督域适应(UDA)新方法。现有对抗域适应方法在处理分类问题中多模态分布的不同域时,难以有效对齐。本文方法具备两大特点:首先,采用ResNet深层结构与特征金字塔网络(FPN)的尺度分离机制,同时捕捉内容与风格特征;其次,结合新型损失函数与精心选择的已有损失函数,以应对自然图像中存在的尺度、噪声和风格偏移等挑战,这些挑战叠加于多模态(多类)分布之上。该组合损失函数不仅提升了目标域上的模型精度与鲁棒性,还加速了训练收敛。所提方案在Office-Home、Office-31和VisDA-2017数据集上优于现有CNN-based方法,在DomainNet数据集上表现相当。

原文摘要 · Abstract (English)

We present a novel approach for unsupervised domain adaptation (UDA) for natural images. A commonly-used objective for UDA schemes is to enhance domain alignment in representation space even if there is a domain shift in the input space. Existing adversarial domain adaptation methods may not effectively align different domains of multimodal distributions associated with classification problems. Our approach has two main features. Firstly, its neural architecture uses the deep structure of ResNet and the effective separation of scales of feature pyramidal network (FPN) to work with both content and style features. Secondly, it uses a combination of a novel loss function and judiciously selected existing loss functions to train the network architecture. This tailored combination is designed to address challenges inherent to natural images, such as scale, noise, and style shifts, that occur on top of a multi-modal (multi-class) distribution. The combined loss function not only enhances model accuracy and robustness on the target domain but also speeds up training convergence. Our proposed UDA scheme generalizes better than state-of-the-art for CNN-based methods on Office-Home, Office-31, and VisDA-2017 datasets and comaparable for DomainNet dataset.

域适应图像识别特征对齐深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。