InfoNCE让对比学习的表示趋向高斯分布,解释了为何这类模型常呈现高斯特性。
InfoNCE Induces Gaussian Distribution
- 通过理论分析证明InfoNCE促使高维表示投影趋于多元高斯分布
- 在宽松假设下,加入微小正则项可实现相同渐近高斯行为
- 适用于理解对比学习机制或设计新模型的研究者
对比学习已成为现代表示学习的核心,使大规模无标签数据可用于任务特定和通用(基础)模型的训练。典型的对比损失是InfoNCE及其变体。本文表明,InfoNCE目标会诱导对比训练中产生的表示呈现出高斯结构。我们在两个互补的设定下建立这一结果:首先,在某些对齐与集中假设下,高维表示的投影渐近趋近于多元高斯分布;其次,在更宽松的假设下,添加一个渐近消失的正则项(促进低特征范数与高特征熵),同样可获得类似渐近结果。我们在合成数据和CIFAR-10数据集上,使用多种编码器架构和规模进行实验,均展示了稳定的高斯行为。该视角为对比表示中普遍观察到的高斯性提供了原则性解释。由此生成的高斯模型使得对学习表示进行严谨分析成为可能,并有望支持广泛的对比学习应用。
原文摘要 · Abstract (English)
Contrastive learning has become a cornerstone of modern representation learning, allowing training with massive unlabeled data for both task-specific and general (foundation) models. A prototypical loss in contrastive training is InfoNCE and its variants. In this work, we show that the InfoNCE objective induces Gaussian structure in representations that emerge from contrastive training. We establish this result in two complementary regimes. First, we show that under certain alignment and concentration assumptions, projections of the high-dimensional representation asymptotically approach a multivariate Gaussian distribution. Next, under less strict assumptions, we show that adding a small asymptotically vanishing regularization term that promotes low feature norm and high feature entropy leads to similar asymptotic results. We support our analysis with experiments on synthetic and CIFAR-10 datasets across multiple encoder architectures and sizes, demonstrating consistent Gaussian behavior. This perspective provides a principled explanation for commonly observed Gaussianity in contrastive representations. The resulting Gaussian model enables principled analytical treatment of learned representations and is expected to support a wide range of applications in contrastive learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。