arXiv:2502.09507cs.LGcs.CV2025-02ICML被引 10

CLIP的泛化能力依赖训练数据多样性,但组合泛化可能弱于领域泛化。

When and How Does CLIP Enable Domain and Compositional Generalization?

  • 用受控数据分布训练CLIP,研究其泛化能力来源。
  • 领域多样性是两种泛化的核心,但组合泛化在特定条件下更弱。
  • 中间层共享表征是成功泛化的关键机制,适合模型可解释性研究者。

对比视觉-语言模型如CLIP的卓越泛化性能常归因于其训练数据分布的多样性。然而,关键问题仍未解决:当在多种领域混合数据上训练时,CLIP能否泛化到完全未见的领域(领域泛化)?能否在部分可见领域中泛化到未见类别(组合泛化)?哪些因素影响这种泛化?为此,我们使用系统构建的、控制领域多样性和物体类别暴露程度的训练分布训练CLIP模型。实验表明,领域多样性对两类泛化均至关重要;然而,当训练分布包含测试领域的一个次优子集时,组合泛化能力可能显著弱于领域泛化。通过数据驱动和机制分析,我们发现成功的泛化依赖于中间层中足够共享的表示学习。

原文摘要 · Abstract (English)

The remarkable generalization performance of contrastive vision-language models like CLIP is often attributed to the diversity of their training distributions. However, key questions remain unanswered: Can CLIP generalize to an entirely unseen domain when trained on a diverse mixture of domains (domain generalization)? Can it generalize to unseen classes within partially seen domains (compositional generalization)? What factors affect such generalization? To answer these questions, we trained CLIP models on systematically constructed training distributions with controlled domain diversity and object class exposure. Our experiments show that domain diversity is essential for both domain and compositional generalization, yet compositional generalization can be surprisingly weaker than domain generalization when the training distribution contains a suboptimal subset of the test domain. Through data-centric and mechanistic analyses, we find that successful generalization requires the learning of sufficiently shared representations in intermediate layers and circuits.

CLIP泛化能力多模态表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。