arXiv:2606.00424cs.AI2026-06被引 2

用弱模型当批评者,让强模型自我改进。

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight

  • 让弱模型只指出改进方向,不直接给答案
  • 通过自教师信号逐步吸收高质量批评
  • 适合资源有限时对大模型进行持续优化

随着大语言模型能力增强,弱监督者可能无法为复杂输出提供可靠标签、偏好或最终判断,限制了从弱到强的泛化和可扩展监督。本文研究一种更可行的弱监督形式:将弱模型作为批评者而非标注者或裁判。弱批评者无需完成任务或选择正确答案,只需提供不误导的修正方向,帮助强模型更好利用自身知识。我们提出渐进式在线策略批评蒸馏(OPCD),通过自适应自教师信号筛选高质量批评,并将其指导行为融入强模型。在推理和对齐基准上的实验表明,该方法能在训练过程中持续提升强模型表现,为弱监督下的可扩展监督提供了有效路径。

原文摘要 · Abstract (English)

As large language models become stronger, weak supervisors may fail to provide reliable labels, preferences, or final judgments for complex outputs, limiting both weak-to-strong generalization and scalable oversight. We study a more tractable form of weak supervision: using a weak model as a critic rather than as a labeler or judge. Instead of solving the task or selecting the correct answer, the weak critic only needs to provide a non-misleading revision direction that helps the strong model better use its own knowledge. We call this setting *weak-critic strong oversight*. We first show that weak critiques can improve frozen strong models at inference time, and that critique quality is key to this improvement. We then propose progressive on-policy critique distillation (**OPCD**), which filters high-quality critiques and distills critic-guided behavior into the strong model through adaptive self-teacher signals. Experiments on reasoning and alignment benchmarks show that our method improves strong models over training epochs, suggesting an effective path for scalable oversight with weak supervision.

强化学习模型蒸馏弱监督大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。