arXiv:2509.24896cs.CV2025-09

用多模态模型提升无源域自适应的标注效率与性能

DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation

  • 融合视觉语言模型与人工标注,形成双路监督信号
  • 在多个基准上超越现有方法,实现新最佳性能
  • 适合关注高效域自适应与少样本学习的研究者

无源主动域自适应(SFADA)通过主动学习选择少量人工标签,将源模型知识迁移至无标签目标域。尽管近期研究引入视觉-语言(ViL)模型以提升伪标签质量或特征对齐,但通常将ViL与数据监督视为独立来源,缺乏有效融合。为此,我们提出双重主动学习与多模态基础模型(DAM),将ViL模型的多模态监督与稀疏人工标注结合,形成双路监督信号。DAM初始化稳定的ViL引导目标,并采用双向蒸馏机制,在迭代适配过程中促进目标模型与双监督信号间的相互知识传递。大量实验表明,DAM在多个SFADA基准和主动学习策略下持续优于现有方法,创下新纪录。

原文摘要 · Abstract (English)

Source-free active domain adaptation (SFADA) enhances knowledge transfer from a source model to an unlabeled target domain using limited manual labels selected via active learning. While recent domain adaptation studies have introduced Vision-and-Language (ViL) models to improve pseudo-label quality or feature alignment, they often treat ViL-based and data supervision as separate sources, lacking effective fusion. To overcome this limitation, we propose Dual Active learning with Multimodal (DAM) foundation model, a novel framework that integrates multimodal supervision from a ViL model to complement sparse human annotations, thereby forming a dual supervisory signal. DAM initializes stable ViL-guided targets and employs a bidirectional distillation mechanism to foster mutual knowledge exchange between the target model and the dual supervisions during iterative adaptation. Extensive experiments demonstrate that DAM consistently outperforms existing methods and sets a new state-of-the-art across multiple SFADA benchmarks and active learning strategies.

域自适应多模态主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。