让弱模型主动指导强模型,提升推理能力
Alice: Proactive Learning with Teacher's Demonstrations for Weak-to-Strong Generalization
- 通过探测教师模型不确定性,主动生成优化示范
- 数学推理提升22.62%,逻辑推理提升12.11%
- 适合大模型监督与知识迁移场景
大型语言模型(LLM)能力增强带来有效人工监管的挑战。弱到强泛化(W2SG)为利用弱模型监督强模型提供了前景。传统W2SG采用被动学习,弱教师提供含噪示范,限制学生模型在训练中发挥自身能力。本文提出Alice框架,通过探测教师模型的不确定性,结合其响应作为示范,引导学生自生成改进后的响应以实现监督。针对教师与学生间能力差距较大的情况,引入级联Alice,采用分层训练:弱教师先监督中间模型,再逐级指导更强模型。实验表明,该方法显著提升W2SG性能,在三项关键任务中表现优异:知识推理提升4.0%,数学推理提升22.62%,逻辑推理提升12.11%。验证了新范式在实现更稳健知识迁移与监督结果方面的有效性。
原文摘要 · Abstract (English)
The growing capabilities of large language models (LLMs) present a key challenge of maintaining effective human oversight. Weak-to-strong generalization (W2SG) offers a promising framework for supervising increasingly capable LLMs using weaker ones. Traditional W2SG methods rely on passive learning, where a weak teacher provides noisy demonstrations to train a strong student. This hinders students from employing their knowledge during training and reaching their full potential. In this work, we introduce Alice (pro{A}ctive {l}earning w{i}th tea{c}her's D{e}monstrations), a framework that leverages complementary knowledge between teacher and student to enhance the learning process. We probe the knowledge base of the teacher model by eliciting their uncertainty, and then use these insights together with teachers' responses as demonstrations to guide student models in self-generating improved responses for supervision. In addition, for situations with significant capability gaps between teacher and student models, we introduce cascade Alice, which employs a hierarchical training approach where weak teachers initially supervise intermediate models, who then guide stronger models in sequence. Experimental results demonstrate that our method significantly enhances the W2SG performance, yielding substantial improvements in three key tasks compared to the original W2SG: knowledge-based reasoning (+4.0%), mathematical reasoning (+22.62%), and logical reasoning (+12.11%). This highlights the effectiveness of our new W2SG paradigm that enables more robust knowledge transfer and supervision outcome.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。