弱教师标签训练强学生,能显著提升模型泛化性能。
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
- 用随机特征岭回归建模师生关系,推导测试误差的确定性等价式。
- 学生误差的缩放律可优于教师,即使教师误差不随样本量下降。
- 在偏差主导和方差主导场景下均实现性能提升,适合数据标注研究者。
机器学习中常使用模型对数据打标,再用这些标签训练更强模型。弱到强泛化现象体现了这种两阶段方法的优势:尽管学生使用弱教师生成的不完美标签进行训练,但其性能仍超过教师。本文研究随机特征岭回归(RFRR)下的师生模型,通过推导学生测试误差的确定性等价式,揭示在某些条件下,学生误差的缩放律优于教师。这一改进既可在偏差主导,也可在方差主导设置中实现。令人惊讶的是,即便教师测试误差不随样本量减少,学生仍可能达到极小最大最优率。
原文摘要 · Abstract (English)
It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenomenon of weak-to-strong generalization exemplifies the advantage of this two-stage procedure: a strong student is trained on imperfect labels obtained from a weak teacher, and yet the strong student outperforms the weak teacher. In this paper, we show that the potential improvement is substantial, in the sense that it affects the scaling law followed by the test error. Specifically, we consider students and teachers trained via random feature ridge regression (RFRR). Our main technical contribution is to derive a deterministic equivalent for the excess test error of the student trained on labels obtained via the teacher. Via this deterministic equivalent, we then identify regimes in which the scaling law of the student improves upon that of the teacher, unveiling that the improvement can be achieved both in bias-dominated and variance-dominated settings. Strikingly, the student may attain the minimax optimal rate regardless of the scaling law of the teacher -- in fact, when the test error of the teacher does not even decay with the sample size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。