提出用数据谱结构优化对比学习,提升训练效率与稳定性。
Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes
- 基于对齐、温度和批数据谱的非渐近界,推导梯度范数上界
- 有效秩衡量各向异性,在ImageNet-100上提速15%且保持精度
- 批内白化可降低梯度方差,理论与实测结果高度吻合
我们推导了非渐近的谱区间,以约束平方InfoNCE梯度范数,其依赖于对齐度、温度参数和批数据谱结构,恢复了$1/τ^{2}$规律,并在合成数据和ImageNet上紧密追踪批均梯度。采用有效秩$R_{\mathrm{eff}}$作为各向异性代理,设计了谱感知的批数据选择策略,包括一种快速贪心构建器。在ImageNet-100上,Greedy-64相比随机采样将达到67.5%准确率的时间减少15%(相比Pool--P3减少24%),在同等精度下;CIFAR-10也表现出类似收益。批内白化促进各向同性,使50步梯度方差降低1.37倍,与理论上限匹配。
原文摘要 · Abstract (English)
We derive non-asymptotic spectral bands that bound the squared InfoNCE gradient norm via alignment, temperature, and batch spectrum, recovering the \(1/τ^{2}\) law and closely tracking batch-mean gradients on synthetic data and ImageNet. Using effective rank \(R_{\mathrm{eff}}\) as an anisotropy proxy, we design spectrum-aware batch selection, including a fast greedy builder. On ImageNet-100, Greedy-64 cuts time-to-67.5\% top-1 by 15\% vs.\ random (24\% vs.\ Pool--P3) at equal accuracy; CIFAR-10 shows similar gains. In-batch whitening promotes isotropy and reduces 50-step gradient variance by \(1.37\times\), matching our theoretical upper bound.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。