用视觉语言模型增强无源域适应,提升迁移性能。
ViLAaD: Enhancing "Attracting and Dispersing'' Source-Free Domain Adaptation with Vision-and-Language Model
- 基于吸引与分散框架,引入视觉语言模型初始化目标域适应。
- 在四个基准上实现最优性能,尤其在强初始化条件下优势明显。
- 方法灵活可扩展,适合多场景无源域适应任务研究者使用。
无源域适应(SFDA)旨在不访问源数据的情况下,将预训练的源模型迁移到不同领域的目标数据集。传统方法受限于源模型编码信息和无标签目标数据。近期虽有利用辅助资源的方法出现,但仍处于初期阶段。本文提出一种新方法,通过在现有SFDA框架中引入视觉-语言(ViL)模型来扩展辅助信息。具体地,我们基于广泛采用的吸引与分散(AaD)方法,将其核心原理推广以自然融合ViL模型作为目标适应的强力初始化。所提方法称为ViL增强型AaD(ViLAaD),在保持AaD简单性和灵活性的同时,显著提升适应性能。实验表明,使用多种ViL模型时,ViLAaD始终优于AaD及零样本分类效果。此外,该方法可无缝集成至交替优化框架中,并结合提示调优与额外目标进行扩展。在四个SFDA基准上的大量实验显示,其改进版本ViLAaD++在闭集、部分集和开集等多种场景下均达到当前最优表现。
原文摘要 · Abstract (English)
Source-Free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to a target dataset from a different domain without access to the source data. Conventional SFDA methods are limited by the information encoded in the pre-trained source model and the unlabeled target data. Recently, approaches leveraging auxiliary resources have emerged, yet remain in their early stages, offering ample opportunities for research. In this work, we propose a novel method that incorporates auxiliary information by extending an existing SFDA framework using Vision-and-Language (ViL) models. Specifically, we build upon Attracting and Dispersing (AaD), a widely adopted SFDA technique, and generalize its core principle to naturally integrate ViL models as a powerful initialization for target adaptation. Our approach, called ViL-enhanced AaD (ViLAaD), preserves the simplicity and flexibility of the AaD framework, while leveraging ViL models to significantly boost adaptation performance. We validate our method through experiments using various ViL models, demonstrating that ViLAaD consistently outperforms both AaD and zero-shot classification by ViL models, especially when both the source model and ViL model provide strong initializations. Moreover, the flexibility of ViLAaD allows it to be seamlessly incorporated into an alternating optimization framework with ViL prompt tuning and extended with additional objectives for target model adaptation. Extensive experiments on four SFDA benchmarks show that this enhanced version, ViLAaD++, achieves state-of-the-art performance across multiple SFDA scenarios, including Closed-set SFDA, Partial-set SFDA, and Open-set SFDA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。