arXiv:2605.02116cs.LG2026-05中稿 · ICML被引 1

解释对比学习为何有效,揭示负样本越多越好的理论原因。

Statistical Consistency and Generalization of Contrastive Representation Learning

  • 构建统一统计理论,证明对比损失可一致收敛于最优排序
  • 推导出风险与检索性能的定量关系,且负样本数增加时泛化界反而变好
  • 适用于理解视觉-语言模型训练机制的研究者和算法工程师

对比表示学习(CRL)是现代基础模型的核心。现有理论存在三大局限:(i)CRL的统计一致性尚不明确;(ii)已有泛化界随负样本数量增加而恶化,与实证优势矛盾;(iii)CRL的检索性能缺乏理论分析。本文建立统一的统计学习理论:针对下游任务,采用基于AUC的总体评估标准,证明对比损失在最优排序下具有统计一致性;进一步建立校准型不等式,量化超额对比风险与超额检索次优性的关系。针对上游训练,分别对监督与自监督对比目标推导出泛化界:分别为 $O(1/m + 1/ oot{n})$ 与 $O(1/ oot{m} + 1/ oot{n})$,其中 $m$ 为负样本数,$n$ 为锚点数。该界解释了大规模负样本的实证优势,并揭示 $m$ 与 $n$ 的显式权衡。大规模视觉-语言模型实验验证了理论预测。

原文摘要 · Abstract (English)

Contrastive representation learning (CRL) underpins many modern foundation models. Despite recent theoretical progress, existing analyses suffer from several key limitations: (i) the statistical consistency of CRL remains poorly understood; (ii) available generalization bounds deteriorate as the number of negative samples increases, contradicting the empirical benefits of large negative sets; and (iii) the retrieval performance of CRL has received limited theoretical attention. In this paper, we develop a unified statistical learning theory for CRL. For downstream tasks, we evaluate retrieval quality using an AUC-type population criterion and show that the contrastive loss is \emph{statistically consistent} with optimal ranking. We further establish a \emph{calibration-style inequality} that quantitatively relates excess contrastive risk to excess retrieval suboptimality. For upstream training, we study both supervised and self-supervised contrastive objectives and derive generalization bounds of order $O(1/m + 1/\sqrt{n})$ and $O(1/\sqrt{m} + 1/\sqrt{n})$, respectively, where $m$ denotes the number of negative samples and $n$ the number of anchor points. These bounds not only explain the empirical advantages of large negative sets but also reveal an explicit trade-off between $m$ and $n$. Extensive experiments on large-scale vision--language models corroborate our theoretical predictions.

对比学习泛化理论统计一致性视觉-语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。