弱教师也能教会强学生,且学生表现更好,无需大模型
Weak-to-Strong Generalization Even in Random Feature Networks, Provably
- 用随机特征网络模拟师生关系,教师单元少,学生单元多
- 学生仅学教师标注数据,仍可超越教师性能
- 早期停止是关键,且该现象有明确量化边界
弱到强泛化(Burns et al., 2024)指强学生(如GPT-4)从弱教师(如GPT-2)学习任务后,表现显著优于教师。本文证明此现象不依赖强学习器。研究采用两层随机特征网络:底部随机固定,顶部可训练。弱教师使用少量单元在总体分布上训练;强学生使用远多于教师的单元,仅在教师生成的标签数据上训练。我们证明并理解学生如何在仅依赖教师标签的情况下超越教师。同时揭示早期停止是实现该现象的关键机制,并给出此模型下弱到强泛化的定量极限。
原文摘要 · Abstract (English)
Weak-to-Strong Generalization (Burns et al., 2024) is the phenomenon whereby a strong student, say GPT-4, learns a task from a weak teacher, say GPT-2, and ends up significantly outperforming the teacher. We show that this phenomenon does not require a strong learner like GPT-4. We consider student and teacher that are random feature models, described by two-layer networks with a random and fixed bottom layer and a trained top layer. A "weak" teacher, with a small number of units (i.e. random features), is trained on the population, and a "strong" student, with a much larger number of units (i.e. random features), is trained only on labels generated by the weak teacher. We demonstrate, prove, and understand how the student can outperform the teacher, even though trained only on data labeled by the teacher. We also explain how such weak-to-strong generalization is enabled by early stopping. Importantly, we also show the quantitative limits of weak-to-strong generalization in this model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。