为深度对比学习的泛化能力提供新理论分析,突破传统方法对网络深度的强依赖。
Generalization Analysis for Deep Contrastive Representation Learning
- 从参数量和权重范数双角度推导泛化界,不依赖对比学习的样本组大小k。
- 新方法避免深度指数增长问题,比现有工作更优且与普通损失学习理论接轨。
- 适用于研究对比学习理论的学者,尤其关注深度神经网络泛化性的研究者。
本文针对深度对比表示学习框架中的无监督风险,提供了泛化边界分析。该框架使用深度神经网络作为表示函数。我们从两个角度展开:一是基于参数量的界,随网络整体规模增长;二是基于权重矩阵范数的界,随权重范数变化。忽略对数因子后,这些界与对比学习中样本组大小k无关。据我们所知,仅有一项工作具备类似性质,但其采用不同证明策略,且对网络深度呈极强指数依赖,源于使用剥皮技术。我们的结果通过利用关于样本上一致范数的覆盖数强大结果,规避了这一问题。此外,我们引入损失增强技术,进一步降低对矩阵范数的依赖及隐含的深度依赖。事实上,本方法可生成多个具有类似结构依赖关系的对比学习界,使其与普通损失函数样本复杂度研究的理论框架相衔接。
原文摘要 · Abstract (English)
In this paper, we present generalization bounds for the unsupervised risk in the Deep Contrastive Representation Learning framework, which employs deep neural networks as representation functions. We approach this problem from two angles. On the one hand, we derive a parameter-counting bound that scales with the overall size of the neural networks. On the other hand, we provide a norm-based bound that scales with the norms of neural networks' weight matrices. Ignoring logarithmic factors, the bounds are independent of $k$, the size of the tuples provided for contrastive learning. To the best of our knowledge, this property is only shared by one other work, which employed a different proof strategy and suffers from very strong exponential dependence on the depth of the network which is due to a use of the peeling technique. Our results circumvent this by leveraging powerful results on covering numbers with respect to uniform norms over samples. In addition, we utilize loss augmentation techniques to further reduce the dependency on matrix norms and the implicit dependence on network depth. In fact, our techniques allow us to produce many bounds for the contrastive learning setting with similar architectural dependencies as in the study of the sample complexity of ordinary loss functions, thereby bridging the gap between the learning theories of contrastive learning and DNNs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。