弱监督强模型为何能超越教师:理论揭示泛化与校准的边界
The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration
- 通过上下界分析,揭示强模型泛化性能受弱模型误差和优化目标制约
- 强模型过度优化会因依赖弱监督信号而损害泛化能力
- 回归任务中证明强模型可至少超越教师其分歧程度的性能提升
弱到强的泛化指弱监督训练的强模型表现优于其弱教师,为对齐超人类模型与人类价值观提供新路径。本文在分类场景下建立强模型泛化误差的上下界,指出主要限制源于弱模型的泛化误差及其优化目标;同时推导出强模型校准误差的上下界。理论表明:弱模型需具备良好泛化能力且预测校准良好;强模型训练需谨慎平衡,过度优化会导致其过度依赖弱监督信号而削弱泛化能力。在回归场景中,将Charikar等(2024)的工作扩展至基于KL散度的损失函数,证明强学生可至少以两者分歧程度的幅度超越弱教师。大量实验验证了理论结论。
原文摘要 · Abstract (English)
Weak-to-strong generalization, where weakly supervised strong models outperform their weaker teachers, offers a promising approach to aligning superhuman models with human values. To deepen the understanding of this approach, we provide theoretical insights into its capabilities and limitations. First, in the classification setting, we establish upper and lower generalization error bounds for the strong model, identifying the primary limitations as stemming from the weak model's generalization error and the optimization objective itself. Additionally, we derive lower and upper bounds on the calibration error of the strong model. These theoretical bounds reveal two critical insights: (1) the weak model should demonstrate strong generalization performance and maintain well-calibrated predictions, and (2) the strong model's training process must strike a careful balance, as excessive optimization could undermine its generalization capability by over-relying on the weak supervision signals. Finally, in the regression setting, we extend the work of Charikar et al. (2024) to a loss function based on Kullback-Leibler (KL) divergence, offering guarantees that the strong student can outperform its weak teacher by at least the magnitude of their disagreement. We conduct sufficient experiments to validate our theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。