arXiv:2603.17474cs.CVcs.AI2026-03

通过注入有益噪声增强注意力机制,提升跨域图像识别的鲁棒性。

Revisiting Cross-Attention Mechanisms: Leveraging Beneficial Noise for Domain-Adaptive Learning

  • 引入有益噪声正则化交叉注意力,聚焦内容而非风格
  • 在VisDA-2017上比CDTrans提升2.3%,卡车类达5.9%增益
  • 适合处理外观和尺度差异大的跨域学习任务

无监督域适应(UDA)旨在将带标签源域知识迁移到无标签目标域,但常受域间差异和尺度变化影响导致性能下降。现有基于交叉注意力的Transformer虽能对齐特征,却难以在显著外观与尺度变化下保持语义一致性。为此,本文提出‘有益噪声’概念,通过可控扰动正则化交叉注意力,引导模型忽略风格干扰、关注内容信息。我们设计了域自适应跨尺度匹配(DACSM)框架,包含域自适应变换器(DAT),用于分离共享内容与特定风格;以及跨尺度匹配(CSM)模块,实现多分辨率特征自适应对齐。DAT在交叉注意力中引入有益噪声,实现渐进式域迁移并增强鲁棒性,生成内容一致、风格不变的表示;CSM保障尺度变化下的语义一致性。在VisDA-2017、Office-Home和DomainNet上的大量实验表明,DACSM达到当前最优性能,相比CDTrans在VisDA-2017上提升最高2.3%;尤其在挑战性的‘truck’类别上获得5.9%提升,验证了有益噪声在处理尺度差异方面的有效性。结果表明,结合域迁移、噪声增强注意力与尺度感知对齐,可有效实现鲁棒的跨域表征学习。

原文摘要 · Abstract (English)

Unsupervised Domain Adaptation (UDA) seeks to transfer knowledge from a labeled source domain to an unlabeled target domain but often suffers from severe domain and scale gaps that degrade performance. Existing cross-attention-based transformers can align features across domains, yet they struggle to preserve content semantics under large appearance and scale variations. To explicitly address these challenges, we introduce the concept of beneficial noise, which regularizes cross-attention by injecting controlled perturbations, encouraging the model to ignore style distractions and focus on content. We propose the Domain-Adaptive Cross-Scale Matching (DACSM) framework, which consists of a Domain-Adaptive Transformer (DAT) for disentangling domain-shared content from domain-specific style, and a Cross-Scale Matching (CSM) module that adaptively aligns features across multiple resolutions. DAT incorporates beneficial noise into cross-attention, enabling progressive domain translation with enhanced robustness, yielding content-consistent and style-invariant representations. Meanwhile, CSM ensures semantic consistency under scale changes. Extensive experiments on VisDA-2017, Office-Home, and DomainNet demonstrate that DACSM achieves state-of-the-art performance, with up to +2.3% improvement over CDTrans on VisDA-2017. Notably, DACSM achieves a +5.9% gain on the challenging "truck" class of VisDA, evidencing the strength of beneficial noise in handling scale discrepancies. These results highlight the effectiveness of combining domain translation, beneficial-noise-enhanced attention, and scale-aware alignment for robust cross-domain representation learning.

域自适应交叉注意力噪声增强图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。