用反向KL提升强模型性能,有效抑制弱模型噪声干扰。
Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
- 改用反向KL替代正向KL,聚焦高置信预测,降低噪声影响。
- 理论证明反向KL可保证强模型超越弱监督者,差距等于两者分歧度。
- 实验显示反向损失在多数场景下让强模型表现优于传统方法。
随着大语言模型逼近超人级表现,确保其与人类价值观和能力对齐变得日益复杂。弱到强泛化通过利用弱模型的预测指导强模型,展现巨大潜力,但其效果常受限于弱预测中的噪声和不准确性。为此,我们提出一种理论驱动的方法:以反向KL散度取代正向KL散度——后者因覆盖全质量分布,可能过拟合至不完美的弱信号。反向KL具有零强制效应,优先关注高置信度预测,有效削弱不可靠弱监督的影响。理论上,我们扩展了现有边界,并推导出正向与反向KL的更紧下界,证明反向KL至少可提供与正向KL相当的保证。特别地,当强模型在最后线性层上微调时,反向KL保证其性能超越弱监督者,差距等于二者分歧程度。实验表明,反向KL与反向交叉熵使强模型在多数设置中成功超越正向KL与标准交叉熵训练的模型,凸显其实际优势。
原文摘要 · Abstract (English)
As large language models advance toward superhuman performance, ensuring their alignment with human values and abilities grows increasingly complex. Weak-to-strong generalization offers a promising approach by leveraging predictions from weaker models to guide stronger systems, but its effectiveness could be constrained by the inherent noise and inaccuracies in these weak predictions. To address this, we propose a theoretically grounded approach that replaces forward KL divergence-whose mass-covering behavior risks overfitting to imperfect weak signals-with reverse KL divergence. Reverse KL divergence's zero-forcing effect prioritizes high-confidence predictions, effectively mitigating the influence of unreliable weak supervision. Theoretically, we extend existing bounds and derive tighter lower bounds for both forward and reverse KL divergence, establishing that reverse KL achieves at least comparable guarantees to forward KL. Notably, when a sufficiently pre-trained strong model is fine-tuned on the last linear layer, reverse KL guarantees that it outperforms its weak supervisor by the magnitude of their disagreement. Empirically, we demonstrate that reverse KL and reverse cross-entropy enable strong models to successfully outperform those trained with forward KL and standard cross-entropy across most settings, highlighting the practical advantages of these reverse losses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。