arXiv:2501.13031cs.LG2025-01

提出一种概率模型,解释非对比学习为何有效。

A Probabilistic Model for Non-Contrastive Learning

  • 构建隐变量统计模型,将自监督学习形式化为最大似然估计。
  • 当数据增强信息丰富时,模型解等价于主成分分析(PCA)。
  • 揭示了非对比损失与生成模型的内在联系,适合理论研究者。

自监督学习(SSL)旨在通过数据增强编码语义相似性,从无标签数据中提取有意义的表示。尽管该方法广受欢迎,但其理论基础仍不充分。例如,目前尚不清楚常用SSL损失函数是否可关联到一个统计模型,如同普通最小二乘、广义线性模型或主成分分析(PCA)自然作为底层生成过程的最大似然估计一样。本文提出一种用于自监督学习的隐变量统计模型,该模型具有一个有趣性质:根据数据增强的信息量,模型的最大似然估计(MLE)要么退化为PCA,要么趋近于一个简单的非对比损失。我们分析了该模型,并通过实验验证了相关发现。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) aims to find meaningful representations from unlabeled data by encoding semantic similarities through data augmentations. Despite its current popularity, theoretical insights about SSL are still scarce. For example, it is not yet known whether commonly used SSL loss functions can be related to a statistical model, much in the same as OLS, generalized linear models or PCA naturally emerge as maximum likelihood estimates of an underlying generative process. In this short paper, we consider a latent variable statistical model for SSL that exhibits an interesting property: Depending on the informativeness of the data augmentations, the MLE of the model either reduces to PCA, or approaches a simple non-contrastive loss. We analyze the model and also empirically illustrate our findings.

自监督学习概率模型非对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。