让视觉Transformer更专注前景,提升跨域适应能力
PCaM: A Progressive Focus Attention-Based Information Fusion Method for Improving Vision Transformer Domain Adaptation
- 分阶段过滤背景信息,聚焦跨域前景语义融合
- 在Office-Home等数据集上显著提升准确率,达新最好效果
- 轻量无侵入,可无缝接入现有ViT域适应框架
无监督域适应(UDA)旨在将带标签源域的知识迁移到无标签目标域。基于视觉变换器(ViTs)的近期方法通过注意力对齐实现了强劲性能,但我们发现一个关键问题:前景物体不匹配,即不同域间前景物体尺寸和空间分布差异导致注意力一致性减弱,阻碍有效对齐。为此,我们提出渐进式聚焦交叉注意力机制(PCaM),在交叉注意力过程中逐步滤除背景信息,使模型能聚焦并融合跨域的判别性前景语义。我们还引入注意力引导损失,显式引导注意力关注任务相关区域,增强跨域注意力一致性。PCaM轻量、架构无关,易于集成到现有基于ViT的UDA流程中。在Office-Home、DomainNet、VisDA-2017及遥感数据集上的大量实验表明,PCaM显著提升适应性能,达到新的最优结果,验证了注意力引导前景融合在域适应中的有效性。
原文摘要 · Abstract (English)
Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Recent UDA methods based on Vision Transformers (ViTs) have achieved strong performance through attention-based feature alignment. However, we identify a key limitation: foreground object mismatch, where the discrepancy in foreground object size and spatial distribution across domains weakens attention consistency and hampers effective domain alignment. To address this issue, we propose the Progressive Focus Cross-Attention Mechanism (PCaM), which progressively filters out background information during cross-attention, allowing the model to focus on and fuse discriminative foreground semantics across domains. We further introduce an attentional guidance loss that explicitly directs attention toward task-relevant regions, enhancing cross-domain attention consistency. PCaM is lightweight, architecture-agnostic, and easy to integrate into existing ViT-based UDA pipelines. Extensive experiments on Office-Home, DomainNet, VisDA-2017, and remote sensing datasets demonstrate that PCaM significantly improves adaptation performance and achieves new state-of-the-art results, validating the effectiveness of attention-guided foreground fusion for domain adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。