arXiv:2510.10980cs.LGcs.CV2025-10

用几何方法证明巴洛孪生模型能实现最优表征效率

On the Optimal Representation Efficiency of Barlow Twins: An Information-Geometric Interpretation

  • 通过费舍尔信息矩阵的谱特性定义表征效率
  • 理论证明巴洛孪生使表征矩阵趋近单位阵,效率达1
  • 为自监督学习提供新几何视角,适合研究者参考

自监督学习在无标签数据下取得显著成功,但缺乏统一理论框架来理解与比较不同范式的学习效率。本文提出一种新颖的信息几何框架,将表征效率η定义为学习表征空间的有效内在维数与环境维数之比,其中有效维数由编码器诱导的统计流形上的费舍尔信息矩阵(FIM)谱性质决定。在此框架下,我们对巴洛孪生方法进行理论分析,在合理假设下证明其可通过将表征的交叉相关矩阵推向单位矩阵,从而诱导各向同性FIM,实现最优表征效率(η=1)。该工作为巴洛孪生的有效性提供了严格理论基础,并为分析自监督学习算法提供了新的几何视角。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has achieved remarkable success by learning meaningful representations without labeled data. However, a unified theoretical framework for understanding and comparing the efficiency of different SSL paradigms remains elusive. In this paper, we introduce a novel information-geometric framework to quantify representation efficiency. We define representation efficiency $η$ as the ratio between the effective intrinsic dimension of the learned representation space and its ambient dimension, where the effective dimension is derived from the spectral properties of the Fisher Information Matrix (FIM) on the statistical manifold induced by the encoder. Within this framework, we present a theoretical analysis of the Barlow Twins method. Under specific but natural assumptions, we prove that Barlow Twins achieves optimal representation efficiency ($η= 1$) by driving the cross-correlation matrix of representations towards the identity matrix, which in turn induces an isotropic FIM. This work provides a rigorous theoretical foundation for understanding the effectiveness of Barlow Twins and offers a new geometric perspective for analyzing SSL algorithms.

自监督学习表征效率信息几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。