arXiv:2601.20154cs.LG2026-01

从谱视角统一解释自监督学习,揭示其核心机理。

Spectral Ghost in Representation Learning: from Component Analysis to Self-Supervised Learning

  • 基于谱分析构建表示学习的理论框架。
  • 揭示现有成功自监督方法的共同谱本质。
  • 为高效算法设计提供可解释的指导,适合研究者与工程师参考。

自监督学习(SSL)通过利用海量无标签数据提升表示学习的性能,在各类下游任务中展现出显著效果。然而,尽管出现了多种不同的学习目标和训练流程,当前仍缺乏清晰统一的理论理解。这种理论缺失阻碍了表示学习的发展,导致算法设计缺乏原则性指导,实际应用也缺乏充分依据。本文提出一种基于谱表示视角的理论分析,揭示了现有成功自监督算法的共性本质,并构建了一个统一的分析框架。该框架不仅有助于深入理解已有方法,还为未来开发更高效、易用的表示学习算法提供了可遵循的理论基础,推动其实用化进程。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has improved empirical performance by unleashing the power of unlabeled data for practical applications. Specifically, SSL extracts the representation from massive unlabeled data, which will be transferred to a plenty of down streaming tasks with limited data. The significant improvement on diverse applications of representation learning has attracted increasing attention, resulting in a variety of dramatically different self-supervised learning objectives for representation extraction, with an assortment of learning procedures, but the lack of a clear and unified understanding. Such an absence hampers the ongoing development of representation learning, leaving a theoretical understanding missing, principles for efficient algorithm design unclear, and the use of representation learning methods in practice unjustified. The urgency for a unified framework is further motivated by the rapid growth in representation learning methods. In this paper, we are therefore compelled to develop a principled foundation of representation learning. We first theoretically investigate the sufficiency of the representation from a spectral representation view, which reveals the spectral essence of the existing successful SSL algorithms and paves the path to a unified framework for understanding and analysis. Such a framework work also inspires the development of more efficient and easy-to-use representation learning algorithms with principled way in real-world applications.

表示学习自监督谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。