arXiv:2605.11870cs.LGcs.IT2026-05

用信息论解释聚类型自监督学习的原理,揭示了蒸馏与中心化背后的理论依据。

Information theoretic underpinning of self-supervised learning by clustering

  • 将自监督学习建模为KL散度优化,通过约束教师分布防止模式坍缩
  • 推导出使用逆聚类先验归一化,等价于常见的批量中心化操作
  • 为现有自监督方法提供理论支撑,指导未来研究方向

自监督学习(SSL)是构建人工智能基础模型的关键工具。近年来,其进展得益于对SSL原则的深入探讨和大量实证研究。本文旨在发展SSL的理论基础,聚焦于深度聚类方法。类比于有监督学习,我们将SSL形式化为KL散度优化问题,并通过在教师分布上施加优化约束来防止模式坍缩。由此导出使用逆聚类先验进行归一化的方法。利用Jensen不等式,我们证明该归一化可简化为广泛采用的批量中心化过程。蒸馏与中心化是当前自监督学习中的常见启发式手段,但本工作首次从理论上为其提供了支撑。所建立的理论模型不仅解释了已有成功方法的有效性,还为未来研究指明了方向。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) is recognized as an essential tool for building foundation models for Artificial Intelligence applications. The advances in SSL have been made thanks to vigorous arguments about the principles of SSL and through extensive empirical research. The aim of this paper is to contribute to the development of the underpinning theory of SSL, focusing on the deep clustering approach. By analogy to supervised learning, we formulate SSL as K-L divergence optimization. The mode collapse is prevented by imposing an optimisation constraint on the teacher distribution. This leads to normalization using inverse cluster priors. We show that using Jensen inequality this normalization simplifies to the popular batch centering procedure. Distillation and centering are common {heuristics-based} practices in SSL, {but our work underpins them theoretically.} The theoretical model developed not only supports specific existing successful SSL methods, but also suggests directions for future investigations.

自监督学习信息论聚类理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。