揭示InfoNCE的聚类机制,提出可调目标的新损失函数
Understanding InfoNCE: Transition Probability Matrix Induced Feature Clustering
- 用转移概率矩阵建模数据增强动态,揭示特征聚类原理
- 新损失函数可调节收敛目标,提升跨领域性能稳定性
- 适合研究对比学习理论或想优化表示学习的工程师
对比学习已成为视觉、语言和图领域无监督表示学习的核心方法,其中InfoNCE是主流目标函数。尽管其在实践中表现优异,但理论基础仍不充分。本文引入显式特征空间以建模样本的增强视图,并使用转移概率矩阵捕捉数据增强过程的动力学特性。我们证明,InfoNCE通过优化两个视图共享同一来源的概率,使其趋近于由该矩阵定义的恒定目标,从而自然诱导出表示空间中的特征聚类。基于此洞察,我们提出新型损失函数SC-InfoNCE,通过引入可调节的收敛目标,灵活控制特征相似性对齐。通过对目标矩阵进行缩放,SC-InfoNCE实现对特征相似性对齐的灵活调控,使训练目标更匹配下游数据的统计特性。在图像、图和文本等基准数据集上的实验表明,SC-InfoNCE在多个领域均表现出一致且可靠的强性能。
原文摘要 · Abstract (English)
Contrastive learning has emerged as a cornerstone of unsupervised representation learning across vision, language, and graph domains, with InfoNCE as its dominant objective. Despite its empirical success, the theoretical underpinnings of InfoNCE remain limited. In this work, we introduce an explicit feature space to model augmented views of samples and a transition probability matrix to capture data augmentation dynamics. We demonstrate that InfoNCE optimizes the probability of two views sharing the same source toward a constant target defined by this matrix, naturally inducing feature clustering in the representation space. Leveraging this insight, we propose Scaled Convergence InfoNCE (SC-InfoNCE), a novel loss function that introduces a tunable convergence target to flexibly control feature similarity alignment. By scaling the target matrix, SC-InfoNCE enables flexible control over feature similarity alignment, allowing the training objective to better match the statistical properties of downstream data. Experiments on benchmark datasets, including image, graph, and text tasks, show that SC-InfoNCE consistently achieves strong and reliable performance across diverse domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。