提出视觉自监督学习三大必要原则,揭示现有方法的理论统一性。
Three Necessary Principles for Self-Supervised Visual Representation Learning

- 从观察、预测、正则三方面构建统一学习框架
- 实验证明缺任一原则都会导致表征退化或失效
- 适用于理解主流自监督模型原理的研究者
我们指出,无标签视觉表征学习需要同时满足三个不重叠目标:跨增强视图的语义不变性、像素级空间预测和表征非退化性。将其形式化为观测、预测与正则化原则。证明:(i)在无负样本对齐下,仅结合观测与预测时恒等编码器是全局最小解;(ii)两个目标在编码器输出层梯度互补且结构不冲突;(iii)动量编码器收敛至与在线编码器相同不动点,但无法保证收敛时不失效。对比对齐仅提供自我限制的崩溃抵抗,通过显式梯度衰减论证。移除预测会主动剥夺空间训练信号;移除观测则破坏跨视图语义一致性;在所研究规模下,任一任务均不可替代第三项。所有主要自监督方法均为单一统一能量分解的特例。每项理论主张均配以受控实验,包括用于评估预测空间影响的补丁检索任务。
原文摘要 · Abstract (English)
We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy. We formalize these as the observation, prediction, and regularization principles and prove (i) that combining observation and prediction without regularization admits the constant encoder as a global minimizer under negative-free alignment; (ii) that the two objectives are gradient-complementary and structurally non-conflicting at the encoder output; and (iii) that the momentum encoder converges to the same fixed point as the online encoder and provides no collapse guarantee at convergence. Contrastive alignment provides only self-limiting collapse resistance, formalized via an explicit gradient-decay argument. Dropping prediction withholds the spatial training signal by construction; dropping observation forfeits cross-view semantic invariance by construction; at the scale we study, no pair substitutes for the third. Every major self-supervised method is a special case of a single unified energy decomposition. We pair every theoretical claim with a controlled experiment, including a patch-retrieval evaluation for the spatial consequence of prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。