arXiv:2603.22908cs.CVcs.LG2026-03

解决黑盒域适应中模型与语言模型的语义鸿沟问题

Adaptive Dual-Teacher Distillation with Subnetwork Rectification for Bridging Semantic Gaps in Black-Box Domain Adaptation

  • 用双教师自适应融合黑盒模型与视觉语言模型预测
  • 通过子网正则化降低噪声标签导致的过拟合
  • 适合无源数据/参数的实用域适应场景

在无法访问源数据或源模型参数的黑盒域适应(BBDA)设定下,迁移知识仅限于黑盒源模型的预测结果。现有方法通过伪标签优化或利用视觉语言模型(ViL)获取知识,但常因黑盒模型的任务特定知识与ViL的语言对齐语义先验之间存在本质差异,导致整合效果不佳、适配性能下降。为此,我们提出自适应双教师蒸馏与子网修正框架(DDSR),显式调和这两类互补但不一致的知识源。DDSR采用自适应预测融合策略,结合黑盒源模型与ViL的输出,生成目标域可靠伪标签;基于子网的正则化机制通过强制输出一致性与梯度发散性,缓解噪声监督带来的过拟合;同时,迭代优化的目标域预测逐步精炼伪标签与ViL提示,提升语义对齐。最终,类别级原型通过自训练进一步优化目标预测。在多个基准数据集上的大量实验表明,DDSR持续优于当前最优方法,包括部分可访问源数据或模型参数的方法。

原文摘要 · Abstract (English)

Assuming that neither source data nor source model parameters are accessible, black-box domain adaptation (BBDA) represents a highly practical yet challenging setting, where transferable knowledge is limited to the predictions of a black-box source model. Existing approaches exploit such knowledge via pseudo-label refinement or by leveraging vision-language models (ViLs), but they often fail to reconcile the inherent discrepancy between task-specific knowledge from black-box models and language-aligned semantic priors of ViLs, resulting in suboptimal integration and degraded adaptation performance. To address this challenge, we propose adaptive Dual-Teacher Distillation with Subnetwork Rectification (DDSR), a framework that explicitly reconciles these complementary yet inconsistent knowledge sources. DDSR employs an adaptive prediction fusion strategy to integrate predictions from the black-box source model and a ViL, generating reliable pseudo-labels for the target domain. A subnetwork-based regularization mechanism mitigates overfitting to noisy supervision by enforcing output consistency and gradient divergency. Furthermore, progressively improved target predictions iteratively refine both pseudo-labels and ViL prompts, enhancing semantic alignment. Finally, class-wise prototypes are used to further optimize target predictions via self-training. Extensive experiments on multiple benchmark datasets demonstrate that DDSR consistently outperforms state-of-the-art methods, including those with access to source data or source model parameters.

域适应黑盒学习视觉语言模型知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。