为对比学习提供可信赖的风险评估,解决传统方法失效问题。
Tight PAC-Bayesian Risk Certificates for Contrastive Learning
- 基于PAC-Bayesian框架,考虑SimCLR中正负样本重用导致的依赖性。
- 在CIFAR-10上验证,对比损失与下游分类风险边界显著更紧。
- 适用于关注模型泛化性、尤其是对比学习理论分析的研究者。
对比表示学习通过数据增强构建语义相似样本对(正例)与独立样本(负例),使前者嵌入更接近。尽管该方法在基础模型中广泛应用且效果显著,其统计理论仍不完善。已有工作虽提出泛化误差界,但多为平凡结果或依赖不切实际的强假设。本文针对主流SimCLR框架,提出非平凡的PAC-Bayesian风险证书,充分考虑其正例被重复用作负例所引发的强依赖性,使经典PAC或PAC-Bayesian界限不再适用。进一步结合数据增强与温度缩放等SimCLR特性,优化下游分类损失的界,并推导出对比零一损失的风险证书。实验表明,该方法在CIFAR-10上的对比损失与预测风险边界均显著优于先前结果。
原文摘要 · Abstract (English)
Contrastive representation learning is a modern paradigm for learning representations of unlabeled data via augmentations -- precisely, contrastive models learn to embed semantically similar pairs of samples (positive pairs) closer than independently drawn samples (negative samples). In spite of its empirical success and widespread use in foundation models, statistical theory for contrastive learning remains less explored. Recent works have developed generalization error bounds for contrastive losses, but the resulting risk certificates are either vacuous (certificates based on Rademacher complexity or $f$-divergence) or require strong assumptions about samples that are unreasonable in practice. The present paper develops non-vacuous PAC-Bayesian risk certificates for contrastive representation learning, considering the practical considerations of the popular SimCLR framework. Notably, we take into account that SimCLR reuses positive pairs of augmented data as negative samples for other data, thereby inducing strong dependence and making classical PAC or PAC-Bayesian bounds inapplicable. We further refine existing bounds on the downstream classification loss by incorporating SimCLR-specific factors, including data augmentation and temperature scaling, and derive risk certificates for the contrastive zero-one risk. The resulting bounds for contrastive loss and downstream prediction are much tighter than those of previous risk certificates, as demonstrated by experiments on CIFAR-10.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。