arXiv:2512.02831stat.MLcs.LG2025-12

揭示对比学习在跨域泛化中的理论局限,提出新评估框架。

Revisiting Theory of Contrastive Learning for Domain Generalization

  • 构建考虑分布偏移与新类别空间的泛化边界
  • 证明对比学习在跨域任务上的性能依赖于分布差异度
  • 为新类别场景提供可证明的性能保证,适合研究者参考

对比学习是自监督表征学习中最流行且强大的方法之一,旨在将语义相似样本在潜在空间中拉近,而使不相似样本分离。现有理论假设下游任务类别来自预训练阶段相同的潜在类别分布,但在真实场景中,下游任务不仅可能在同一标签空间内出现分布偏移,还可能引入新或更广的标签空间,带来领域泛化挑战。本文提出新的泛化边界,明确考虑两类失配:领域偏移和领域泛化。具体分析两种情形:(i) 下游任务从相同潜在类别空间中采样但分布发生偏移;(ii) 涉及预训练中未见的新标签空间。分析表明,对比学习表征性能取决于预训练与下游分布间的统计差异。该扩展视角使我们能够对超出预训练类别集的平均分类任务,给出学习表征性能的可证明保障。

原文摘要 · Abstract (English)

Contrastive learning is among the most popular and powerful approaches for self-supervised representation learning, where the goal is to map semantically similar samples close together while separating dissimilar ones in the latent space. Existing theoretical methods assume that downstream task classes are drawn from the same latent class distribution used during the pretraining phase. However, in real-world settings, downstream tasks may not only exhibit distributional shifts within the same label space but also introduce new or broader label spaces, leading to domain generalization challenges. In this work, we introduce novel generalization bounds that explicitly account for both types of mismatch: domain shift and domain generalization. Specifically, we analyze scenarios where downstream tasks either (i) draw classes from the same latent class space but with shifted distributions, or (ii) involve new label spaces beyond those seen during pretraining. Our analysis reveals how the performance of contrastively learned representations depends on the statistical discrepancy between pretraining and downstream distributions. This extended perspective allows us to derive provable guarantees on the performance of learned representations on average classification tasks involving class distributions outside the pretraining latent class set.

对比学习领域泛化理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。