arXiv:2510.21808cs.CVcs.AI2025-10

提升CLIP在跨域零样本学习中的性能,解决知识迁移与模态对齐难题

Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning

  • 引入语义关系结构损失,增强类别间知识迁移效率
  • 在I2AwA和I2WebV上达到新最优,显著超越现有方法
  • 适合关注跨域零样本学习与视觉语言模型优化的研究者

数据标注成本高昂推动了数据受限场景下深度学习模型的研究。现有范式难以平衡跨域迁移与跨类别泛化,催生了领域自适应零样本学习(DAZSL)的需求。尽管视觉-语言模型(如CLIP)在该领域具备天然优势,但现有研究未充分挖掘其潜力。将CLIP用于DAZSL面临两大核心挑战:缺乏语义关系引导导致跨类别知识迁移低效,目标域微调中跨模态对齐性能下降。为此,我们提出语义关系增强型CLIP适配器(SRE-CLIP),融合语义关系结构损失与跨模态对齐保留策略。作为首个基于CLIP的DAZSL方法,SRE-CLIP在I2AwA与I2WebV基准上实现最先进性能,显著优于现有方法。

原文摘要 · Abstract (English)

The high cost of data annotation has spurred research on training deep learning models in data-limited scenarios. Existing paradigms, however, fail to balance cross-domain transfer and cross-category generalization, giving rise to the demand for Domain-Adaptive Zero-Shot Learning (DAZSL). Although vision-language models (e.g., CLIP) have inherent advantages in the DAZSL field, current studies do not fully exploit their potential. Applying CLIP to DAZSL faces two core challenges: inefficient cross-category knowledge transfer due to the lack of semantic relation guidance, and degraded cross-modal alignment during target domain fine-tuning. To address these issues, we propose a Semantic Relation-Enhanced CLIP (SRE-CLIP) Adapter framework, integrating a Semantic Relation Structure Loss and a Cross-Modal Alignment Retention Strategy. As the first CLIP-based DAZSL method, SRE-CLIP achieves state-of-the-art performance on the I2AwA and I2WebV benchmarks, significantly outperforming existing approaches.

零样本学习CLIP跨域适应视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。