arXiv:2510.23484cs.LGcs.CG2025-10NeurIPS被引 2

用最小生成树长度正则化,提升自监督学习的表征质量。

T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning

  • 基于表示点间最小生成树长度设计正则项
  • 有效防止特征维数坍缩并提升分布均匀性
  • 适用于图像等数据的自监督表征学习

自监督学习(SSL)通过强制输入变换(如旋转、模糊)下的不变性来学习无标签数据的表征。近期研究指出,优质表征需满足两个关键特性:(i) 避免维度坍缩——即特征仅占据低维子空间;(ii) 提升诱导分布的均匀性。本文提出T-REGS,一种基于最小生成树(MST)长度的简单正则化框架。理论分析表明,T-REGS在任意紧致黎曼流形上可同时缓解维度坍缩并促进分布均匀性。在合成数据及经典SSL基准上的实验验证了该方法在提升表征质量方面的有效性。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has emerged as a powerful paradigm for learning representations without labeled data, often by enforcing invariance to input transformations such as rotations or blurring. Recent studies have highlighted two pivotal properties for effective representations: (i) avoiding dimensional collapse-where the learned features occupy only a low-dimensional subspace, and (ii) enhancing uniformity of the induced distribution. In this work, we introduce T-REGS, a simple regularization framework for SSL based on the length of the Minimum Spanning Tree (MST) over the learned representation. We provide theoretical analysis demonstrating that T-REGS simultaneously mitigates dimensional collapse and promotes distribution uniformity on arbitrary compact Riemannian manifolds. Several experiments on synthetic data and on classical SSL benchmarks validate the effectiveness of our approach at enhancing representation quality.

自监督学习表征学习正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。