arXiv:2507.01028cs.LGcs.AI2025-07被引 3

解析自监督学习中两种稳定策略的原理,揭示其防坍缩机制。

Dual Perspectives on Non-Contrastive Self-Supervised Learning

  • 从优化与动力系统双视角分析非对比学习中的稳定策略。
  • 线性情况下,无稳定策略必致表示坍缩,有则可保持渐近稳定。
  • 理论结合实验,适用于研究自监督学习机制的学者。

非对比自监督学习中常用的梯度截断(stop gradient)与指数移动平均(exponential moving average)迭代方法,虽不直接优化原始目标或任意光滑函数,但能有效避免表示坍缩,实践中表现优异。本文从优化与动力系统双重角度进行分析。基于Tian21的工作,但无需其额外假设,我们证明在线性情形下,若不使用梯度截断或指数移动平均,则最小化原始目标函数必然导致坍缩。反之,我们明确刻画了这两种方法对应动力系统的平衡点为参数空间中的代数簇,并证明其一般为渐近稳定。理论结果通过真实与合成数据的实验得到验证。

原文摘要 · Abstract (English)

The {\em stop gradient} and {\em exponential moving average} iterative procedures are commonly used in non-contrastive approaches to self-supervised learning to avoid representation collapse, with excellent performance in downstream applications in practice. This presentation investigates these procedures from the dual viewpoints of optimization and dynamical systems. We show that, in general, although they {\em do not} optimize the original objective, or {\em any} other smooth function, they {\em do} avoid collapse Following~\citet{Tian21}, but without any of the extra assumptions used in their proofs, we then show using a dynamical system perspective that, in the linear case, minimizing the original objective function without the use of a stop gradient or exponential moving average {\em always} leads to collapse. Conversely, we characterize explicitly the equilibria of the dynamical systems associated with these two procedures in this linear setting as algebraic varieties in their parameter space, and show that they are, in general, {\em asymptotically stable}. Our theoretical findings are illustrated by empirical experiments with real and synthetic data.

自监督学习表示坍缩动力系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。