破解非独立同分布下对比学习的泛化难题,揭示样本需求与表征复杂度的关系。
Generalization Analysis for Supervised Contrastive Representation Learning under Non-IID Settings
- 基于U统计量思想,建立非独立同分布下的对比学习泛化界
- 每类样本数仅需对数级增长,与该类表征空间的覆盖数相关
- 适用于线性映射和神经网络,为实际数据复用提供理论支撑
对比表示学习近年来在多个领域取得显著成功,但其泛化行为的理论理解仍不充分。现有研究多假设用于对比学习的数据元组独立同分布,然而实践中常受限于有限的可重用标注数据,不得不在元组间重复使用数据以构建足够大的数据集,导致先前假设失效。本文首次在非独立同分布设置下对对比学习框架进行泛化分析,更贴近实际场景。受U统计量文献启发,我们推导出泛化界,表明每类所需样本数量与该类可学习特征表示空间的覆盖数的对数成正比。进一步将主结果应用于常见函数类(如线性映射和神经网络),得到相应的过拟合风险上界。
原文摘要 · Abstract (English)
Contrastive Representation Learning (CRL) has achieved impressive success in various domains in recent years. Nevertheless, the theoretical understanding of the generalization behavior of CRL has remained limited. Moreover, to the best of our knowledge, the current literature only analyzes generalization bounds under the assumption that the data tuples used for contrastive learning are independently and identically distributed. However, in practice, we are often limited to a fixed pool of reusable labeled data points, making it inevitable to recycle data across tuples to create sufficiently large datasets. Therefore, the tuple-wise independence condition imposed by previous works is invalidated. In this paper, we provide a generalization analysis for the CRL framework under non-$i.i.d.$ settings that adheres to practice more realistically. Drawing inspiration from the literature on U-statistics, we derive generalization bounds which indicate that the required number of samples in each class scales as the logarithm of the covering number of the class of learnable feature representations associated to that class. Next, we apply our main results to derive excess risk bounds for common function classes such as linear maps and neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。