arXiv:2510.24812cs.LGstat.ML2025-10NeurIPS被引 5

从线性模型到非线性模型,揭示弱监督强模型泛化机制

From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning

  • 通过梯度下降分析线性CNN到两层ReLU CNN的迁移过程
  • 数据稀缺时依赖良性过拟合,数据充足时早期通过标签修正提升性能
  • 首次在非线性框架下严格证明弱到强泛化,适合机器学习理论研究者

弱到强泛化指由较弱模型监督训练的更强模型能超越其教师。现有理论多局限于抽象框架或线性/随机特征模型。本文首次对从线性卷积网络(弱)到两层ReLU卷积网络(强)的弱到强泛化进行形式化分析。考虑包含依赖标签的信号与独立噪声的结构化数据,分析强模型在弱模型标注数据上进行梯度下降的动态过程。根据信号-噪声比将数据分为两类:数据稀缺与数据充裕。在数据稀缺情形,泛化由良性过拟合实现或因有害过拟合失败,我们刻画了二者间的临界边界;在数据充裕情形,泛化在训练早期通过标签修正出现,但过度训练会导致性能下降。

原文摘要 · Abstract (English)

Weak-to-strong generalization refers to the phenomenon where a stronger model trained under supervision from a weaker one can outperform its teacher. While prior studies aim to explain this effect, most theoretical insights are limited to abstract frameworks or linear/random feature models. In this paper, we provide a formal analysis of weak-to-strong generalization from a linear CNN (weak) to a two-layer ReLU CNN (strong). We consider structured data composed of label-dependent signals of varying difficulty and label-independent noise, and analyze gradient descent dynamics when the strong model is trained on data labeled by the pretrained weak model. Our analysis identifies two regimes -- data-scarce and data-abundant -- based on the signal-to-noise characteristics of the dataset, and reveals distinct mechanisms of weak-to-strong generalization. In the data-scarce regime, generalization occurs via benign overfitting or fails via harmful overfitting, depending on the amount of data, and we characterize the transition boundary. In the data-abundant regime, generalization emerges in the early phase through label correction, but we observe that overtraining can subsequently degrade performance.

泛化理论非线性模型特征学习深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。