arXiv:2606.26617cs.LG2026-06

首次建立对比学习的缩略线性模型理论,揭示其规模扩展规律。

Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling

论文配图:Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling
图 1 · 摘自论文原文
  • 基于配对高斯潜变量,研究缩略视图下的双线性对比评分训练
  • 导出风险分解公式,明确优化与样本噪声如何随模型和数据变化
  • 提供模型、数据与计算量平衡的理论指导,适合关注对比学习机制的研究者

规模定律描述了学习性能随模型大小、数据规模和计算资源的变化规律。尽管已有理论工作建立了缩略线性回归的规模定律,但对比表示学习的相关理论仍不完善。本文在配对高斯潜变量设定下,研究一种用于对比学习的缩略线性模型。学习者仅观测两个相关变量的缩略视图,通过全批量经验梯度下降训练双线性对比得分。我们在对齐幂律谱和对比源条件下分析高斯负二次对比代理函数,导出风险分解为不可约风险、近似误差、梯度下降偏差、梯度下降方差及交叉项。交叉项由偏差和方差控制,因此不影响上界规模。主定理给出了关于缩略维度 $M$、样本数 $N$ 及有效优化时长 $L_{\mathrm{eff}}γ$ 的显式规模定律。与标准线性回归相比,对比设置需学习两视图间交互,这改变了优化与有限样本噪声随模型规模、数据量和训练时间的缩放方式。该工作为理解对比学习的规模行为迈出第一步,并为模型、数据与优化计算的平衡提供理论指引。

原文摘要 · Abstract (English)

Scaling laws describe how learning performance varies with model size, data size, and compute. While recent theoretical work has established scaling laws for sketched linear regression, much less is understood for contrastive representation learning. In this paper, we study a sketched linear model for contrastive learning under a paired Gaussian latent-variable setup. The learner observes only sketched views of two correlated variables and trains a bilinear contrastive score by full-batch empirical gradient descent. We analyze a Gaussian-negative quadratic contrastive surrogate under aligned power-law spectra and a contrastive source condition, where we derive a risk decomposition into irreducible risk, approximation error, GD bias, GD variance, and a cross term. The cross term is controlled by the bias and variance and therefore does not affect the upper-bound scaling. Our main theorem gives an explicit scaling law with respect to sketch dimension $M$, sample size $N$, and effective optimization horizon $L_{\mathrm{eff}}γ$. Compared with standard linear-regression scaling laws, the contrastive setting must learn interactions between two views, and this changes how optimization and finite-sample noise scale with model size, data, and training time. This provides a first theoretical step toward understanding scaling behavior in contrastive learning and gives guidance for balancing model size, data, and optimization compute.

对比学习规模定律理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。