arXiv:2502.08134cs.CV2025-02综述被引 2

对比学习中正负样本设计直接影响模型效果,这篇综述系统梳理了优化策略。

A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters

  • 构建正负样本对的分类体系,归纳现有数据筛选方法
  • 指出高质量样本对可提升表征能力并加快收敛速度
  • 适合关注自监督学习优化的算法研究者阅读

视觉对比学习通过对比相似(正例)与不相似(负例)的数据样本对来学习表征。正负样本对的设计显著影响表征质量、训练效率和计算成本。精心策划的样本对能带来更强的表征能力和更快的收敛速度。随着对比预训练在下游任务中广泛应用,数据筛选变得至关重要。本文尝试建立现有正负样本筛选技术的分类体系,并详细描述各类方法。

原文摘要 · Abstract (English)

Visual contrastive learning aims to learn representations by contrasting similar (positive) and dissimilar (negative) pairs of data samples. The design of these pairs significantly impacts representation quality, training efficiency, and computational cost. A well-curated set of pairs leads to stronger representations and faster convergence. As contrastive pre-training sees wider adoption for solving downstream tasks, data curation becomes essential for optimizing its effectiveness. In this survey, we attempt to create a taxonomy of existing techniques for positive and negative pair curation in contrastive learning, and describe them in detail.

对比学习数据筛选自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。