预训练让强模型超越弱监督者,是其泛化能力跃升的关键
On the Blessing of Pre-training in Weak-to-Strong Generalization

- 用高维单指数模型分析预训练的作用,将其视为几何初始化
- 预训练使模型进入有效区域,实现性能提升并受弱监督偏差制约
- 实证发现W2SG是预训练进程中的相变现象,非模型固有属性
弱到强泛化(W2SG)范式认为,预训练的强模型可超越其弱监督者,但预训练的关键作用在理论和实证上仍不清晰。本文将预训练识别为W2SG出现的必要前提。理论上,在使用尖峰高斯数据的高维单指数模型框架下,将预训练建模为谱初始化步骤;基于此前关于随机初始化下学习失败的不可能性结果,证明当预训练提供几何热启动,使模型进入由扰动强凸几何特征化的有效区域时,W2SG可实现。在此区域内,推导出严格的泛化界,自然捕捉优化动态:初始性能提升后因弱监督者偏差而饱和。实验上,首先通过受控合成模拟验证所有假设与理论洞见;最后通过对大规模语言模型数百个中间预训练检查点的评估,表明W2SG并非内在能力,而是随预训练进程紧密耦合的相变现象。
原文摘要 · Abstract (English)
The paradigm of Weak-to-Strong Generalization (W2SG) suggests that a pre-trained strong model can surpass its weak supervisor, yet the decisive role of pre-training remains theoretically and empirically under-explored. In this work, we identify pre-training as the essential prerequisite for the emergence of W2SG. Theoretically, we formalize the W2SG problem within a high-dimensional single-index model framework using spiked Gaussian data, modeling pre-training as a spectral initialization step. Building upon prior impossibility results regarding the failure of learning under random initialization, we prove that W2SG is achievable when pre-training provides a geometric warm start that places the model within an "effective region" characterized by a perturbed strong-convexity geometry. Within this region, we derive a rigorous generalization bound that naturally captures the optimization dynamics: an initial performance improvement followed by a saturation bottleneck dictated by the weak supervisor's bias. Empirically, we first validate all our assumptions and theoretical insights through controlled synthetic simulations. Finally, through a massive-scale evaluation of hundreds of intermediate pre-training checkpoints from large language models, we demonstrate that W2SG is not an innate capability but emerges via a phase transition tightly coupled with the progression of pre-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。