arXiv:2501.01317cs.LGcs.AI2025-01中稿 · ICLR被引 1

难样本会损害无监督对比学习性能,移除它们反而提升效果。

Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective

  • 构建相似性理论框架,分析样本对间关系
  • 实验证明移除难样本能提升分类准确率
  • 适合关注对比学习机理与性能优化的研究者

无监督对比学习近年表现显著,常接近甚至超越有监督学习。但其学习机制与有监督学习本质不同。以往研究发现,有监督中关键的难样本(决策边界附近样本)在无监督设置中贡献极小。本文令人意外地发现,直接移除难样本虽减少样本量,却能提升对比学习的下游分类性能。为此,我们建立理论框架,建模不同样本对间的相似性,理论分析表明:难样本的存在会损害对比学习的泛化能力。进一步证明,移除难样本、调整边界间隔和温度系数等方法可改善泛化界,从而提升性能。我们提出一种简单高效的难样本筛选机制,并通过实验验证了上述方法的有效性,证实了理论框架的可靠性。

原文摘要 · Abstract (English)

Unsupervised contrastive learning has shown significant performance improvements in recent years, often approaching or even rivaling supervised learning in various tasks. However, its learning mechanism is fundamentally different from supervised learning. Previous works have shown that difficult examples (well-recognized in supervised learning as examples around the decision boundary), which are essential in supervised learning, contribute minimally in unsupervised settings. In this paper, perhaps surprisingly, we find that the direct removal of difficult examples, although reduces the sample size, can boost the downstream classification performance of contrastive learning. To uncover the reasons behind this, we develop a theoretical framework modeling the similarity between different pairs of samples. Guided by this framework, we conduct a thorough theoretical analysis revealing that the presence of difficult examples negatively affects the generalization of contrastive learning. Furthermore, we demonstrate that the removal of these examples, and techniques such as margin tuning and temperature scaling can enhance its generalization bounds, thereby improving performance. Empirically, we propose a simple and efficient mechanism for selecting difficult examples and validate the effectiveness of the aforementioned methods, which substantiates the reliability of our proposed theoretical framework.

对比学习泛化能力难样本理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。