用几何正则化防止表示坍缩,提升自监督学习稳定性与性能
UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures

- 以均匀可展性为正则化目标,通过卡尔森型平方函数实现局部低维流形建模
- 在ImageNet-10上准确率达91.41%,比LeJEPA提升0.83个百分点且方差降低30%
- 适合关注表征质量、模型稳定性的自监督学习研究者
联合嵌入预测架构(JEPAs)训练中的核心难题是防止表示坍缩。LeJEPA通过引入简化的各向同性高斯正则化(SIGReg)来约束嵌入分布,但该目标与流形假设存在矛盾——后者期望嵌入集中于高维空间的低维子集。本文提出UR-JEPA,其目标为小尺度下具有局部切线维数n的均匀n-可展测度,由高斯核平滑的卡尔森型平方函数ℒ^CGLT实现,并辅以琼斯β-数形式。在ImageNet-10上,UR-JEPA(ℒ^CGLT)达到0.9141±0.0014的准确率,较LeJEPA(ℒ^SIGReg)提升+0.83个百分点,且种子标准差降低约30%;在Galaxy10~SDSS、ImageNet-100单种子、EuroSAT三种子任务中,两者收敛精度相近,但UR-JEPA保持更低方差特性。在EuroSAT上,其域内对的准确率达96.0%至96.1%,仅需25倍更小的骨干网络即可媲美大规模遥感基础模型。可视化显示:四数据集上,UR-JEPA的投影输出全局主成分谱在前20~25维处下降4~5个数量级(总维度D=32),而LeJEPA谱近似平坦(首尾比≤3.6)。两方法的逐维边缘分布均为近高斯(谢泼德-沃尔克检验均值W∈[0.992,0.996]),符合迪亚孔尼斯-弗里德曼定理。在相同精度下,二者生成的表示结构迥异。
原文摘要 · Abstract (English)
A central difficulty in training Joint-Embedding Predictive Architectures (JEPAs) is preventing representation collapse. LeJEPA addresses this by enforcing an isotropic Gaussian target on the embeddings via Sketched Isotropic Gaussian Regularization (SIGReg). This target is in tension with the manifold hypothesis, which expects embeddings to concentrate on a low-dimensional subset of the ambient space. We propose \emph{UR-JEPA}, which targets a uniformly $n$-rectifiable measure of local tangent dimension $n$ at small scales, realized through a Gaussian-kernel smoothed Carleson-type square function $\mathcal{L}^{\text{CGLT}}$, with a complementary Jones $β$-number formulation. On Inet10, UR-JEPA($\mathcal{L}^{\text{CGLT}}$) attains $0.9141 \pm 0.0014$ for a $+0.83$\,pp gain over LeJEPA($\mathcal{L}^{\text{SIGReg}}$) with $\sim 30\%$ lower seed standard deviation; on matched-recipe Galaxy10~SDSS, a single-seed ImageNet-$100$ run, and a $3$-seed EuroSAT remote-sensing run, the two methods lie in the same peak-accuracy band at convergence, with UR-JEPA retaining its lower-seed-variance signature. On EuroSAT the in-domain pair is competitive at $96.0$ to $96.1\%$ with large remote-sensing foundation-model transfer at a $25\times$ smaller backbone. The distinction is geometric: direct visualization of the projector output distribution shows that on all four datasets UR--JEPA($\mathcal{L}^{\text{CGLT}}$) produces a global PCA spectrum with a $4$ to $5$ order-of-magnitude drop at index $\sim 20$ to $25$ out of $D = 32$, while LeJEPA's spectrum is near-flat (top-to-bottom ratio at most $3.6$). Per-dimension marginals are simultaneously near-Gaussian for both methods (mean Shapiro-Wilk $W \in [0.992, 0.996]$) as a Diaconis-Freedman consequence. At matched accuracy the two regularizers therefore yield structurally distinct projected representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。