arXiv:2411.01767cs.LG2024-11被引 4

用核理论找出自监督学习中最优数据增强方式

A Theoretical Characterization of Optimal Data Augmentations in Self-Supervised Learning

  • 基于核理论推导出使预训练后表征达标的增强方法
  • 发现增强无需相似于数据或多样化也能有效
  • 适合设计新领域自监督学习的增强策略

数据增强在自监督学习(SSL)的成功中起关键作用。传统观点认为增强应编码不同视图间的不变性,但这一理解忽略了预训练架构的影响,并暗示需要与数据相似且多样的增强。然而,这与实证结果不符,亟需更深入的理论指导以实现增强的合理设计。为此,本文利用核理论,推导了在非对比和对比损失(包括VICReg、Barlow Twins 和谱对比损失)下,能获得目标表征的增强的解析表达式,并提出构造此类增强的算法。分析表明,增强无需与数据相似,也无需多样,且模型架构对最优增强有显著影响。

原文摘要 · Abstract (English)

Data augmentations play an important role in the recent success of self-supervised learning (SSL). While augmentations are commonly understood to encode invariances between different views into the learned representations, this interpretation overlooks the impact of the pretraining architecture and suggests that SSL would require diverse augmentations which resemble the data to work well. However, these assumptions do not align with empirical evidence, encouraging further theoretical understanding to guide the principled design of augmentations in new domains. To this end, we use kernel theory to derive analytical expressions for data augmentations that achieve desired target representations after pretraining. We consider non-contrastive and contrastive losses, namely VICReg, Barlow Twins and the Spectral Contrastive Loss, and provide an algorithm to construct such augmentations. Our analysis shows that augmentations need not be similar to the data to learn useful representations, nor be diverse, and that the architecture has a significant impact on the optimal augmentations.

自监督学习数据增强核理论表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。