弱教师标签也能训练出更强学生模型,原因有三:正则化补偿、结构对齐和特征互补。
On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
- 通过岭回归分析,发现学生可弥补教师正则不足
- 学生若正则结构更匹配任务,性能可超越教师
- 学生能学教师易学特征,自建难学特征以胜过教师
弱到强泛化现象指学生模型在弱教师生成的不完美标签上训练,却仍优于教师。本文通过简单模型的理论分析,揭示三种核心机制:首先,基于岭回归分析,证明学生可通过正则化补偿教师的欠正则化,实现更低测试误差;其次,通过加权岭回归分析,显示学生若具有与目标更契合的正则化结构,可超越教师;第三,在非线性多指标设定下,证明学生可从教师处学习易获取的任务特异性特征,并利用自身更广预训练,捕捉教师无法捕捉的难学特征。
原文摘要 · Abstract (English)
Weak-to-strong generalization, where a student model trained on imperfect labels generated by a weaker teacher nonetheless surpasses that teacher, has been widely observed but the mechanisms that enable it have remained poorly understood. In this paper, through a theoretical analysis of simple models, we uncover three core mechanisms that can drive this phenomenon. First, by analyzing ridge regression, we study the interplay between the teacher and student regularization and prove that a student can compensate for a teacher's under-regularization and achieve lower test error. We also analyze the role of the parameterization regime of the models. Second, by analyzing weighted ridge regression, we show that a student model with a regularization structure more aligned to the target, can outperform its teacher. Third, in a nonlinear multi-index setting, we demonstrate that a student can learn easy, task-specific features from the teacher while leveraging its own broader pre-training to learn hard-to-learn features that the teacher cannot capture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。