让大模型通过多轮辩论自我改进,显著提升推理与一致性。
Self-Improvement of Language Models by Post-Training on Multi-Agent Debate
- 用多智能体辩论生成更优训练信号,强化模型推理能力。
- 在MATH上提升26.87%,数学问答准确率提高21.51%。
- 适合追求模型自进化与强推理能力的研究者使用。
自我改进——模型在无外部监督下超越当前表现——仍是难题,核心在于难以获得强于模型自身输出的训练信号。已有研究显示,多数投票可通过聚合多个样本提供此类信号,缓解语言模型推理中的不一致问题。本文提出多智能体辩论机制,即模型在多轮中协作交换推理过程,比单轮多数投票产生更丰富的信号。我们引入多智能体共识对齐(MACA),采用强化学习后训练模型以有效利用辩论信号。结果表明,基于完整推理轨迹的偏好学习,能更好区分多数与少数推理路径,优于二元共识奖励或监督微调方法。该方法带来三方面提升:模型在多智能体辩论中表现更优(MATH上+26.87%)、独立推理更准确(MathQA上+21.51%)、自我一致性增强(GSM8K上+27.6%)。在未见基准上也展现强泛化能力(GPQA上+16.3%,CommonsenseQA上+11.6%)。
原文摘要 · Abstract (English)
Self-improvement, where models improve beyond their current performance without external supervision, remains a challenge. The core difficulty is sourcing a training signal stronger than what the model itself can currently produce. Majority voting has been shown to provide such a signal by aggregating over multiple samples, helping mitigate some of the inconsistencies in LM reasoning. In this work, we show that multi-agent debate--where models collaborate and exchange reasoning over multiple rounds--provides an even richer signal than single-round majority voting. We introduce Multi-Agent Consensus Alignment (MACA), which uses reinforcement learning (RL) to post-train models to effectively utilize multi-agent debate. We find that preference learning over full reasoning traces, learning to differentiate between majority and minority reasoning, is more effective than binary consensus rewards or SFT-based approaches for leveraging these debate signals. This produces three key improvements: models are (1) better at utilizing the multi-agent debate setting (+26.87% on MATH), (2) individually more accurate (+21.51% on MathQA), and (3) more self-consistent (+27.6% on GSM8K). We also see strong generalization to unseen benchmarks (+16.3% on GPQA, +11.6% on CommonsenseQA).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。