用共识机制动态融合视觉模型与多模态模型,解决无源域适应中的错误信息遗忘问题。
COSMO: Consensus-Driven Shift Modulation for Source-Free Domain Adaptation

- 基于样本级可靠性分配,构建共享共识来协调双模型自适应
- 在四个基准上达到当前最优性能,显著减少对源模型错误证据的依赖
- 适合需要保护源数据隐私的工业级迁移学习场景
无源域适应(SFDA)在不访问源数据的情况下,将源训练模型适配到未标注目标域,适用于隐私或存储受限场景。然而,在显著域偏移下,其自生成的监督信号可能强化源域偏差。预训练视觉语言模型(VLM)可提供互补语义知识,但源模型与VLM在不同目标样本上的可靠性存在差异。现有跨模型引导方法未显式考虑此差异,冲突时可能覆盖有效的源模型证据,导致我们称之为‘源衍生证据遗忘’的问题。本文将VLM引导的SFDA建模为样本级可靠性分配问题,提出共识驱动的域偏移调制(COSMO)。COSMO以锚定共享共识替代专家间指导,先形成偏好更集中预测的样本特定初始共识;在适应过程中,重新聚合两分支演化证据,并根据共识不确定性与训练进度调节最终共识相对于初始锚点的移动范围。该机制保持监督锚定但具备自适应性。在四个基准测试中,使用相同VLM主干,COSMO实现当前最优表现。进一步分析表明,其更好平衡了有效源模型证据保留与互补VLM证据吸收。
原文摘要 · Abstract (English)
Source-free domain adaptation (SFDA) adapts a source-trained model to an unlabeled target domain without source data, a practical setting under privacy or storage constraints. Yet its self-generated supervision can reinforce source bias under substantial domain shifts. Pretrained vision-language models (VLMs) offer complementary semantic knowledge, but the relative reliability of the source model and VLM varies across target samples. Existing cross-model guidance does not explicitly account for this variation and may overwrite valid source-derived evidence under conflict, a failure we term source-derived evidence forgetting. We formulate VLM-guided SFDA as a sample-wise reliability-allocation problem and propose Consensus-Driven Shift Modulation (COSMO). COSMO replaces expert-to-expert guidance with co-adaptation through an anchored shared consensus. It first forms a sample-specific initial consensus that favors the more concentrated prediction. During adaptation, COSMO re-aggregates both branches' evolving evidence and regulates how far the resulting consensus moves from its initial anchor based on consensus uncertainty and training progress. This keeps the shared supervision anchored yet adaptive. Across four benchmarks, COSMO achieves state-of-the-art performance under matched VLM backbones. Further analyses indicate that it better balances the retention of valid source-derived evidence with the absorption of complementary VLM evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。