用多模型集体投票提升大模型推理能力,无需人工标注。
Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback
- 多模型通过投票共进化,用集体一致性优化推理能力。
- 在4个数学推理基准上平均准确率提升16.72%。
- 适合追求高可靠推理的模型集成与自训练场景。
强化学习(RL)显著提升了大语言模型(LLM)的推理能力,但其对昂贵的人工标注数据或复杂奖励模型的依赖严重制约了可扩展性。现有自反馈方法受限于单一模型能力,易导致错误答案的过度自信、奖励欺骗甚至训练崩溃。为此,我们提出基于共进化集体反馈的强化学习框架(RLCCF),实现无外部监督下的多模型协同进化。RLCCF通过最大化集体一致性(CC)来优化模型群体,利用多模型对输出的投票提供奖励信号,且每个模型的投票权重由其自一致性(SC)得分决定,确保更自信的模型贡献更大。得益于多模型输出分布的多样性与互补性,该框架使模型群体在共进化中持续提升推理能力。在四个主流开源大模型及四个数学推理基准上的实验表明,该框架带来显著性能提升,平均准确率相对提高16.72%。值得注意的是,不仅单个模型性能增强,群体多数投票准确率也提升4.51%,证明其能有效拓展模型群体的集体能力边界。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), but its reliance on expensive human-labeled data or complex reward models severely limits scalability. While existing self-feedback methods aim to address this problem, they are constrained by the capabilities of a single model, which can lead to overconfidence in incorrect answers, reward hacking, and even training collapse. To this end, we propose Reinforcement Learning from Coevolutionary Collective Feedback (RLCCF), a novel RL framework that enables multi-model collaborative evolution without external supervision. Specifically, RLCCF optimizes the ability of a model collective by maximizing its Collective Consistency (CC), which jointly trains a diverse ensemble of LLMs and provides reward signals by voting on collective outputs. Moreover, each model's vote is weighted by its Self-Consistency (SC) score, ensuring that more confident models contribute more to the collective decision. Benefiting from the diverse output distributions and complementary abilities of multiple LLMs, RLCCF enables the model collective to continuously enhance its reasoning ability through coevolution. Experiments on four mainstream open-source LLMs across four mathematical reasoning benchmarks demonstrate that our framework yields significant performance gains, achieving an average relative improvement of 16.72\% in accuracy. Notably, RLCCF not only improves the performance of individual models but also enhances the group's majority-voting accuracy by 4.51\%, demonstrating its ability to extend the collective capability boundary of the model collective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。