arXiv:2511.03114cs.LGcs.AI2025-11JMLR被引 7

提出增强重叠理论,解释对比学习为何有效并给出无监督评估指标。

An Augmentation Overlap Theory of Contrastive Learning

  • 基于增强重叠假设推导下游性能的紧致边界。
  • 发现剧烈增强使同类样本支持区域重叠,从而实现聚类。
  • 新评估指标无需额外模块,与下游性能高度一致。

近期自监督对比学习在多个任务中取得显著成功,但其内在工作机制仍不明确。本文首先在广泛采用的条件独立假设下给出了最紧的边界;随后放松该假设,引入更贴近实际的增强重叠假设,并推导出下游性能的渐近闭合边界。所提出的增强重叠理论基于关键洞察:在剧烈数据增强下,同一类样本的支持区域会趋于重叠,因此仅对齐正样本(同一样本的不同增强视图)即可使同类样本自然聚类。此外,从新的增强重叠视角出发,我们开发了一种无监督表示评估指标,几乎不依赖额外模块,却能与下游性能高度一致。代码已公开于 https://github.com/PKU-ML/GARC。

原文摘要 · Abstract (English)

Recently, self-supervised contrastive learning has achieved great success on various tasks. However, its underlying working mechanism is yet unclear. In this paper, we first provide the tightest bounds based on the widely adopted assumption of conditional independence. Further, we relax the conditional independence assumption to a more practical assumption of augmentation overlap and derive the asymptotically closed bounds for the downstream performance. Our proposed augmentation overlap theory hinges on the insight that the support of different intra-class samples will become more overlapped under aggressive data augmentations, thus simply aligning the positive samples (augmented views of the same sample) could make contrastive learning cluster intra-class samples together. Moreover, from the newly derived augmentation overlap perspective, we develop an unsupervised metric for the representation evaluation of contrastive learning, which aligns well with the downstream performance almost without relying on additional modules. Code is available at https://github.com/PKU-ML/GARC.

对比学习理论分析无监督评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。