arXiv:2412.01926cs.LG2024-12被引 5

提出高阶冗余度量,让自监督学习更高效。

Beyond Pairwise Correlations: Higher-Order Redundancies in Self-Supervised Representation Learning

  • 用高阶依赖关系衡量特征空间冗余,突破传统成对相关局限。
  • 实证发现顶尖方法隐式降低冗余,新方法SSLPM性能媲美主流模型。
  • 适合研究表示学习机制、优化自监督模型的学者参考。

多种自监督学习(SSL)方法表明,减少特征嵌入空间中的冗余是有效表示学习的重要手段。然而,这些方法仅关注特征间的成对相关性,忽略了更复杂的冗余形式。为此,本文正式定义了嵌入空间冗余的概念,并引入可捕捉更复杂高阶依赖关系的冗余度量。我们数学分析了这些度量之间的关系,并在常见SSL方法的嵌入空间中进行了实证测量。基于发现,我们提出自监督学习预测性最小化(SSLPM),该方法通过编码器与预测器之间的竞争机制,分别抑制和利用依赖关系以降低冗余。实验表明,SSLPM在性能上可与当前最先进方法比肩,且最佳表现的SSL方法均表现出低嵌入空间冗余,暗示即便无显式冗余减少机制的方法也隐式实现了冗余降低。

原文摘要 · Abstract (English)

Several self-supervised learning (SSL) approaches have shown that redundancy reduction in the feature embedding space is an effective tool for representation learning. However, these methods consider a narrow notion of redundancy, focusing on pairwise correlations between features. To address this limitation, we formalize the notion of embedding space redundancy and introduce redundancy measures that capture more complex, higher-order dependencies. We mathematically analyze the relationships between these metrics, and empirically measure these redundancies in the embedding spaces of common SSL methods. Based on our findings, we propose Self Supervised Learning with Predictability Minimization (SSLPM) as a method for reducing redundancy in the embedding space. SSLPM combines an encoder network with a predictor engaging in a competitive game of reducing and exploiting dependencies respectively. We demonstrate that SSLPM is competitive with state-of-the-art methods and find that the best performing SSL methods exhibit low embedding space redundancy, suggesting that even methods without explicit redundancy reduction mechanisms perform redundancy reduction implicitly.

自监督学习冗余减少特征表示高阶依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。