arXiv:2606.17603cs.LG2026-06被引 1

提出更稳定的球面正则化方法,提升自监督学习的收敛性和性能。

Expanding SPHERE-JEPA: A Family of Statistical Regularizers for the Hypersphere

论文配图:Expanding SPHERE-JEPA: A Family of Statistical Regularizers for the Hypersphere
图 1 · 摘自论文原文
  • 用解析积分替代随机投影,消除梯度噪声,实现确定性优化。
  • 在ImageNet和Galaxy10上验证,收敛更快,性能优于传统切片方法。
  • 不同统计检验影响特征空间结构,适用于不同类型的图像任务。

在自监督学习中,通过强制单位超球面上的表示分布均匀可有效防止表征坍缩。然而,现有框架通常依赖于切片统计正则化(如LeJEPA中的SIGReg和SPHERE-JEPA中的SUSReg),通过沿随机一维方向进行蒙特卡洛采样近似连续目标。这种随机性会引入投影方差,导致训练梯度不稳定,阻碍收敛。本文首次证明,对这些随机投影进行解析积分可天然得到确定性的最大均值差异(MMD),从而绕过切片方法的方差问题。基于此等价性,我们直接在球面上构建了针对MMD、核斯坦因差异(KSD)和相对熵(KL)的全维正则化目标,以强制分布均匀。为避免空间偏差,使用谱理论构造旋转不变核,系统评估两类典型核:平滑指数衰减(热核)与严格频带截断(带限)滤波器。实验表明,去除投影噪声后,优化更稳定,收敛更快,且在ImageNet和Galaxy10上持续优于随机切片正则化。此外,发现统计检验的选择影响潜在空间几何结构:MMD与KSD倾向于局部聚类组织,适合物体中心域;而基于连续KDE的KL则促进细粒度实例分离,在非聚类程序纹理检索任务中表现最优。

原文摘要 · Abstract (English)

In Self-Supervised Learning (SSL), preventing representation collapse by explicitly enforcing a uniform distribution on the unit hypersphere has proven to be effective. However, current frameworks typically rely on sliced statistical regularizers such as SIGReg (used in LeJEPA) and SUSReg (used in SPHERE-JEPA), which approximate this continuous objective via Monte Carlo sampling along random 1D directions. This stochasticity injects projection variance into the training gradients, destabilizing optimization, and hindering convergence. In this work, we first show that analytically integrating out these random projections natively yields a deterministic Maximum Mean Discrepancy (MMD), bypassing the variance of sliced methods. Motivated by this equivalence, we formulate full-dimensional objectives for MMD, Kernel Stein Discrepancy (KSD), and Kullback-Leibler (KL) divergence directly on the sphere to enforce a uniform distribution. To prevent spatial bias, we equip these tests with rotationally invariant kernels constructed via spectral theory, systematically evaluating two canonical families: smooth exponential decay (Heat) and strict frequency cutoff (Bandlimited) filters. Empirically, removing projection-induced noise results in more stable optimization, faster convergence, and consistent improvements over stochastic sliced regularizers on ImageNet and Galaxy10. Furthermore, we reveal that the choice of the statistical test shapes the geometry of the learned latent space: MMD and KSD favor locally clustered organization suitable for object-centric domains, whereas the continuous KDE-based KL divergence promotes fine-grained instance separation, yielding the strongest results on unclustered procedural texture retrieval.

自监督学习球面正则化表示学习优化稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。