解释对比学习为何能提取有效特征,揭示其理论基础。
A Statistical Theory of Contrastive Learning via Approximate Sufficient Statistics
- 用近似充分统计量构建对比学习的统一理论框架。
- 证明最小化对比损失可得到近似充分的编码器。
- 适用于回归与分类任务,适合研究表示学习的学者。
对比学习通过训练模型区分相似与不相似样本,从无标签数据中提取有用表示,推动了基础模型的发展。本文为基于数据增强的对比学习建立新的理论框架,以 SimCLR 为例。方法基于‘近似充分统计量’概念,将其扩展至等价形式和一般 f-散度,证明最小化 SimCLR 等对比损失可获得近似充分的编码器。进一步表明,这些近似充分编码器在下游回归与分类任务中表现良好,性能取决于编码器的充分性以及数据增强引入的误差。线性回归和主题分类的实例展示了结果的广泛适用性。
原文摘要 · Abstract (English)
Contrastive learning -- a modern approach to extract useful representations from unlabeled data by training models to distinguish similar samples from dissimilar ones -- has driven significant progress in foundation models. In this work, we develop a new theoretical framework for analyzing data augmentation-based contrastive learning, with a focus on SimCLR as a representative example. Our approach is based on the concept of \emph{approximate sufficient statistics}, which we extend beyond its original definition in \cite{oko2025statistical} for contrastive language-image pretraining (CLIP) using KL-divergence. We generalize it to equivalent forms and general f-divergences, and show that minimizing SimCLR and other contrastive losses yields encoders that are approximately sufficient. Furthermore, we demonstrate that these near-sufficient encoders can be effectively adapted to downstream regression and classification tasks, with performance depending on their sufficiency and the error induced by data augmentation in contrastive learning. Concrete examples in linear regression and topic classification are provided to illustrate the broad applicability of our results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。