arXiv:2505.22196cs.LG2025-05ICML被引 4

首次建立考虑数据增强的对比学习理论,揭示其对模型误差的直接影响。

An Augmentation-Aware Theory for Self-Supervised Contrastive Learning

  • 提出新的误差界,显式包含数据增强带来的权衡效应
  • 证明特定增强方法能降低理论误差上限
  • 适合关注对比学习机理与增强策略设计的研究者

自监督对比学习已成为从无标签数据中学习有意义表征的强大工具。尽管其实验成功推动了大量理论研究,但现有理论对数据增强的作用仍关注不足,尤其缺乏对具体增强类型影响的分析。本文首次提出一种考虑数据增强的误差界,表明监督风险不仅受无监督风险影响,还受数据增强引发的权衡机制直接约束。在一种新的语义标签假设下,我们进一步分析了特定增强方法如何影响该误差界。最后,通过像素级与表示级实验验证了理论结果的有效性。

原文摘要 · Abstract (English)

Self-supervised contrastive learning has emerged as a powerful tool in machine learning and computer vision to learn meaningful representations from unlabeled data. Meanwhile, its empirical success has encouraged many theoretical studies to reveal the learning mechanisms. However, in the existing theoretical research, the role of data augmentation is still under-exploited, especially the effects of specific augmentation types. To fill in the blank, we for the first time propose an augmentation-aware error bound for self-supervised contrastive learning, showing that the supervised risk is bounded not only by the unsupervised risk, but also explicitly by a trade-off induced by data augmentation. Then, under a novel semantic label assumption, we discuss how certain augmentation methods affect the error bound. Lastly, we conduct both pixel- and representation-level experiments to verify our proposed theoretical results.

对比学习自监督数据增强理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。