arXiv:2511.14166cs.CLcs.AI2025-11被引 4

提出选择性弱到强泛化框架,提升超人类模型对弱监督的鲁棒性。

Selective Weak-to-Strong Generalization

  • 用二分类器判断强模型能否回答问题,仅在必要时使用弱标签
  • 在三个基准上超越基线,自动生成标签提升对齐效果
  • 可跨任务跨难度泛化,适合超对齐场景的模型训练

未来超人类模型将超越人类能力,人类只能提供弱监督。现有弱到强泛化(W2SG)方法通过弱监督微调强模型以实现超越弱监督的泛化,但持续使用弱标签导致鲁棒性下降,部分弱标签反而有害。本文提出选择性弱到强泛化(Selective W2SG)框架,避免不必要的弱监督。训练一个二分类器 P(IK) 判断强模型是否能回答特定问题,并用其自生成标签进行对齐。进一步通过图平滑方法优化弱标签。在三个基准上的实验表明,该方法持续优于竞争基线。进一步分析显示,P(IK) 可跨任务和难度泛化,说明选择性 W2SG 有助于超对齐。

原文摘要 · Abstract (English)

Future superhuman models will surpass the ability of humans and humans will only be able to \textit{weakly} supervise superhuman models. To alleviate the issue of lacking high-quality data for model alignment, some works on weak-to-strong generalization (W2SG) finetune a strong pretrained model with a weak supervisor so that it can generalize beyond weak supervision. However, the invariable use of weak supervision in existing methods exposes issues in robustness, with a proportion of weak labels proving harmful to models. In this paper, we propose a selective W2SG framework to avoid using weak supervision when unnecessary. We train a binary classifier P(IK) to identify questions that a strong model can answer and use its self-generated labels for alignment. We further refine weak labels with a graph smoothing method. Extensive experiments on three benchmarks show that our method consistently outperforms competitive baselines. Further analyses show that P(IK) can generalize across tasks and difficulties, which indicates selective W2SG can help superalignment.

弱监督模型对齐超对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。