用辩论提升弱模型对强模型的监督能力,实现更可靠的对齐。
Debate Helps Weak-to-Strong Generalization
- 让弱模型通过辩论从强模型中提取可信信息
- 弱模型集成可增强对长论辩的利用,提升监督可靠性
- 适用于未来人类无法直接评估超智能模型的对齐场景
当前对已具备能力的模型进行行为对齐,依赖人类提供监督。但未来超人类模型将超越人类能力,导致人类只能提供弱监督。这种监督不足会削弱AI系统安全性。可扩展监督与弱到强泛化是两种互补方法。本文尝试结合二者:利用强预训练模型增强弱人类监督,再以改进后的弱监督训练强模型。我们通过实验验证:小弱模型在强模型辅助下微调,再用弱模型生成的标签微调强模型,发现辩论能帮助弱模型从不可信的强模型中提取可信信息,作为训练上下文。同时,弱模型集成可更好利用强模型生成的长论辩,获得更稳健的监督估计。在OpenAI弱到强NLP基准上的大量实验表明,该组合方法显著提升对齐效果,证明辩论具有促进弱到强泛化的潜力。
原文摘要 · Abstract (English)
Common methods for aligning already-capable models with desired behavior rely on the ability of humans to provide supervision. However, future superhuman models will surpass the capability of humans. Therefore, humans will only be able to weakly supervise superhuman models. This expected deficiency of human evaluation would weaken the safety of future AI systems. Scalable oversight and weak-to-strong generalization are two complementary approaches to tackle this issue. In this paper, we attempt to combine the strengths of these two approaches to further improve alignment. Specifically, we investigate ways of improving human supervision with a strong pretrained model and then supervise the strong model with enhanced weak human supervision. To make iterative empirical progress, we consider an analogy: can we use a strong model to improve weak model supervision and then use it to supervise the strong model? We empirically test it by finetuning a small weak model on ground truth labels with the additional help from a large strong model, and then finetuning the strong model on labels generated by the weak model. We find that debate can assist a weak model in extracting trustworthy information from an untrustworthy strong model, which provides leverage as context on samples when training a weak model. We also show that an ensemble of weak models helps exploit long arguments generated by strong model debaters and obtain a more robust supervision estimate. Extensive experiments on the OpenAI weak-to-strong NLP benchmarks show that the combination approach leads to better alignment, which indicates that debate has the potential to help weak-to-strong generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。