arXiv:2510.21332cs.LGcs.AI2025-10NeurIPS被引 6

弱模型指导强模型时,分布偏移会导致性能下降,新方法可自动选择可信监督者。

Weak-to-Strong Generalization under Distribution Shifts

  • 动态学习弱模型组合权重,提升强模型鲁棒性
  • 跨分布任务性能优于基线超30%,原分布表现相当或更优
  • 能自动识别并赋予更准确弱模型更高权重

随着未来超人类模型日益复杂,人工精确监督其行为可能超出人类能力。已有研究显示,弱模型可有效指导强模型,即弱到强泛化现象。但我们在分布偏移下发现,简单弱到强泛化会显著降低强模型性能。为此,提出RAVEN框架,通过动态学习弱模型组合与强模型参数来增强鲁棒性。在图像分类、文本分类和偏好对齐任务中验证,RAVEN在分布外任务上性能超越基线超过30%,且在分布内任务上表现持平或更优。结果表明,RAVEN能为更准确的弱模型分配更高权重,实现可信监督自动选择。

原文摘要 · Abstract (English)

As future superhuman models become increasingly complex, accurately supervising their behavior may exceed human capabilities. Recent works have demonstrated that in such scenarios, weak models can effectively supervise strong models, a phenomenon known as weak-to-strong generalization. However, we find that naive weak-to-strong generalization fails under distribution shifts, often leading to worse performance of the strong model than its weak supervisors. To address this, we propose RAVEN, a robust weak-to-strong generalization framework that dynamically learns the optimal combinations of weak models in addition to parameters of the strong model. We demonstrate the effectiveness of RAVEN on image classification, text classification, and preference alignment tasks. RAVEN outperforms alternative baselines by over 30% on out-of-distribution tasks while matching or surpassing existing methods on in-distribution tasks. Moreover, our results show that RAVEN assigns higher weights to more accurate weak models, demonstrating its ability to automatically identify trustworthy supervision.

弱监督泛化能力分布外自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。