arXiv:2505.12477cs.LGcs.AI2025-05NeurIPS被引 33

对比自监督学习中重建与联合嵌入的优劣,揭示为何后者在复杂数据上更有效。

Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning

论文配图:Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
图 1 · 摘自论文原文
  • 通过闭式解分析两种方法的表征机制,揭示其对数据增强的依赖差异。
  • 发现当无关特征幅值大时,联合嵌入只需更弱的对齐条件即可最优。
  • 为实际应用中选择方法提供理论依据,适合研究自监督学习机制者阅读。

重建和联合嵌入是自监督学习(SSL)中的两大主流范式。重建方法试图从输入空间的不同视图中恢复原始样本,而联合嵌入方法则在潜在空间中对齐不同视图的表示。两者各有优势,但实践中缺乏明确的选择依据。本文通过闭式解精确刻画了两种方法如何受视图生成过程(如数据增强)影响。结果表明,在非监督学习中,两者均需在增强与无关特征间存在最小对齐才能随样本量增加达到渐近最优。当无关特征幅度较大时,联合嵌入所需的对齐条件严格弱于重建方法,因此更优。该发现不仅厘清了两者的权衡关系,也解释了联合嵌入在真实复杂数据集上的成功经验。

原文摘要 · Abstract (English)

Reconstruction and joint embedding have emerged as two leading paradigms in Self Supervised Learning (SSL). Reconstruction methods focus on recovering the original sample from a different view in input space. On the other hand, joint embedding methods align the representations of different views in latent space. Both approaches offer compelling advantages, yet practitioners lack clear guidelines for choosing between them. In this work, we unveil the core mechanisms that distinguish each paradigm. By leveraging closed form solutions for both approaches, we precisely characterize how the view generation process, e.g. data augmentation, impacts the learned representations. We then demonstrate that, unlike supervised learning, both SSL paradigms require a minimal alignment between augmentations and irrelevant features to achieve asymptotic optimality with increasing sample size. Our findings indicate that in scenarios where these irrelevant features have a large magnitude, joint embedding methods are preferable because they impose a strictly weaker alignment condition compared to reconstruction based methods. These results not only clarify the trade offs between the two paradigms but also substantiate the empirical success of joint embedding approaches on real world challenging datasets.

自监督学习表征学习联合嵌入理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。