从偏差-方差视角揭示强模型在弱监督盲区中的错误风险
Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective

- 引入偏差-方差-协方差框架分析弱到强对齐的失败机制
- 发现强模型方差是导致盲区误判的最显著因素
- 提供早期预警信号,适合改进对齐训练的工程师
弱到强对齐为可扩展监督提供了前景,但在弱模型盲区中,强模型可能自信出错。仅看整体准确率不足以理解此类失败,因为错误不仅取决于强弱模型是否分歧,还与置信度分布有关。本文通过偏差-方差-协方差视角,将误配理论与后训练流程关联,推导出弱到强群体风险的上界,并利用连续置信分数实证分解各成分。在PKU-SafeRLHF和HH-RLHF数据集上评估了四种对齐管道(SFT、RLHF、RLAIF)。引入盲区欺骗度量,分离出强模型自信错误而弱模型不确定的情况,发现强模型方差与该现象最强相关;协方差提供较弱补充信息,表明弱强依赖重要但无法单独解释失败。结果表明,强模型方差可作为早期预警信号,盲区评估能区分失败源于弱监督还是弱模型不确定性区域。
原文摘要 · Abstract (English)
Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak model's blind spots. Understanding such failures requires going beyond aggregate accuracy, since weak-to-strong errors depend not only on whether the strong model disagrees with the weak model, but also on how confidence and uncertainty are distributed across examples. In this work, we analyze weak-to-strong alignment through a bias--variance--covariance lens that connects misfit theory to practical post-training pipelines. We derive a misfit-based upper bound on weak-to-strong population risk and study its empirical components using continuous confidence scores. We evaluate four weak-to-strong pipelines spanning supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), and reinforcement learning from AI feedback (RLAIF) on the PKU-SafeRLHF and HH-RLHF datasets. Using a blind-spot deception metric that isolates cases where the strong model is confidently wrong while the weak model is uncertain, we find that strong-model variance is the quantity most strongly associated with blind-spot deception among the BVC quantities we study. Covariance provides additional but weaker information, indicating that weak--strong dependence matters, but does not by itself explain the observed failures. These results suggest that strong-model variance can serve as an early-warning signal for weak-to-strong deception, while blind-spot evaluation helps distinguish whether failures are inherited from weak supervision or arise in regions of weak-model uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。