从互信息最大化出发,统一解释自监督学习的架构设计原理。
Self-Supervised Representation Learning as Mutual Information Maximization
- 基于变分互信息下界,推导出两种训练范式:自蒸馏与联合优化。
- 揭示停梯度、预测网络等组件在理论上不可替代的必要性。
- 为现有自监督方法提供理论依据,适合研究者深入理解模型设计。
自监督表示学习(SSRL)虽在实践中表现卓越,但其内在机制仍不清晰。现有工作尝试通过信息论目标或启发式规则统一不同方法,但诸如预测网络、停梯度操作和统计正则化等结构元素常被视为经验性添加。本文从第一性原理出发,探究学习目标如何决定优化策略与模型设计。基于变分互信息下界,我们推导出两种训练范式:自蒸馏互信息(SDMI)与联合互信息(JMI),分别对应交替优化与对称联合优化。SDMI天然要求交替优化,使停梯度操作成为理论必需;而JMI可通过对称架构实现联合优化,无需此类组件。在该框架下,SDMI中的预测网络与JMI中的统计正则化均成为互信息目标的可计算代理。我们证明,众多现有SSRL方法均为这两种范式的具体实例或近似。本研究为现有方法的架构选择提供了超越经验便利性的理论解释。
原文摘要 · Abstract (English)
Self-supervised representation learning (SSRL) has demonstrated remarkable empirical success, yet its underlying principles remain insufficiently understood. While recent works attempt to unify SSRL methods by examining their information-theoretic objectives or summarizing their heuristics for preventing representation collapse, architectural elements like the predictor network, stop-gradient operation, and statistical regularizer are often viewed as empirically motivated additions. In this paper, we adopt a first-principles approach and investigate whether the learning objective of an SSRL algorithm dictates its possible optimization strategies and model design choices. In particular, by starting from a variational mutual information (MI) lower bound, we derive two training paradigms, namely Self-Distillation MI (SDMI) and Joint MI (JMI), each imposing distinct structural constraints and covering a set of existing SSRL algorithms. SDMI inherently requires alternating optimization, making stop-gradient operations theoretically essential. In contrast, JMI admits joint optimization through symmetric architectures without such components. Under the proposed formulation, predictor networks in SDMI and statistical regularizers in JMI emerge as tractable surrogates for the MI objective. We show that many existing SSRL methods are specific instances or approximations of these two paradigms. This paper provides a theoretical explanation behind the choices of different architectural components of existing SSRL methods, beyond heuristic conveniences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。