arXiv:2510.10572cs.LG2025-10中稿 · TMLR 2025

用监督学习视角解释对比学习,揭示其内在机制与改进方向。

Understanding Self-supervised Contrastive Learning through Supervised Objectives

  • 将自监督学习建模为监督学习的近似,推导出与InfoNCE相关的损失函数。
  • 提出原型表示偏差和平衡对比损失,解释并提升模型行为。
  • 验证正负样本交互平衡对性能的关键作用,适合研究对比学习者。

自监督表示学习在实践中取得了显著成功,但其理论理解仍不充分。本文从理论角度出发,将自监督表示学习建模为监督表示学习目标的近似。基于此,我们推导出一个与流行的对比损失(如InfoNCE)密切相关的损失函数,揭示了其内在原理。该推导自然引入了原型表示偏差和平衡对比损失的概念,有助于解释并改进自监督学习算法的行为。我们进一步展示了理论框架中的各组件如何对应于对比学习中的既有实践。最后,我们通过实证验证了平衡正负样本交互的影响。所有理论证明均见附录,代码包含在补充材料中。

原文摘要 · Abstract (English)

Self-supervised representation learning has achieved impressive empirical success, yet its theoretical understanding remains limited. In this work, we provide a theoretical perspective by formulating self-supervised representation learning as an approximation to supervised representation learning objectives. Based on this formulation, we derive a loss function closely related to popular contrastive losses such as InfoNCE, offering insight into their underlying principles. Our derivation naturally introduces the concepts of prototype representation bias and a balanced contrastive loss, which help explain and improve the behavior of self-supervised learning algorithms. We further show how components of our theoretical framework correspond to established practices in contrastive learning. Finally, we empirically validate the effect of balancing positive and negative pair interactions. All theoretical proofs are provided in the appendix, and our code is included in the supplementary material.

对比学习自监督理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。