arXiv:2508.21769cs.CVcs.LG2025-08被引 1

让CLIP在未知数据上表现更好,通过分离分类与领域特征。

Domain Generalization in-the-Wild: Disentangling Classification from Domain-Aware Representations

  • 用增强领域感知的分支提升基础模型泛化能力
  • 在33个外部数据集上显著提升对未知领域的适应性
  • 适合研究大模型跨域泛化与鲁棒性优化的学者

评估像CLIP这样的基础模型的域泛化(DG)性能颇具挑战,因为网络规模预训练数据可能已覆盖多个现有基准。因此当前的DG评估既不够具有挑战性,也未能充分测试真正未见数据场景。为更准确评估CLIP在真实世界域泛化中的表现,我们采用两种方法:(1) 在微调后图像网(ImageNet)上的33个多样化数据集上评估,并量化其分布外(OOD)程度;(2) 使用去学习技术使CLIP‘遗忘’某些领域作为近似。结果发现,随着数据集越偏离训练分布,CLIP性能显著下降。为此,我们提出CLIP-DCA(域感知表示解耦分类)。该方法基于观察:标准域不变损失虽旨在使表示域不变,但可能损害基础模型,因其强制丢弃对泛化有益的领域感知特征。我们假设增强领域感知是实现有效域不变分类的前提。CLIP-DCA通过独立的领域头和合成多样领域数据,在编码器中识别并增强领域感知,同时通过解耦机制促进域不变分类。相比现有方法,该模型在更具挑战性的评估中表现显著提升,尤其在更远离训练分布的数据集上。

原文摘要 · Abstract (English)

Evaluating domain generalization (DG) for foundational models like CLIP is challenging, as web-scale pretraining data potentially covers many existing benchmarks. Consequently, current DG evaluation may neither be sufficiently challenging nor adequately test genuinely unseen data scenarios. To better assess the performance of CLIP on DG in-the-wild, a scenario where CLIP encounters challenging unseen data, we consider two approaches: (1) evaluating on 33 diverse datasets with quantified out-of-distribution (OOD) scores after fine-tuning CLIP on ImageNet, and (2) using unlearning to make CLIP `forget' some domains as an approximation. We observe that CLIP's performance deteriorates significantly on more OOD datasets. To address this, we present CLIP-DCA (Disentangling Classification from enhanced domain Aware representations). Our approach is motivated by the observation that while standard domain invariance losses aim to make representations domain-invariant, this can be harmful to foundation models by forcing the discarding of domain-aware representations beneficial for generalization. We instead hypothesize that enhancing domain awareness is a prerequisite for effective domain-invariant classification in foundation models. CLIP-DCA identifies and enhances domain awareness within CLIP's encoders using a separate domain head and synthetically generated diverse domain data. Simultaneously, it encourages domain-invariant classification through disentanglement from the domain features. CLIP-DCA shows significant improvements within this challenging evaluation compared to existing methods, particularly on datasets that are more OOD.

域泛化CLIP解耦表示大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。