arXiv:2508.09670cs.AI2025-08AAAI被引 14

通过多专家互学提升大模型推理能力,解决奖励稀疏问题。

MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement

  • 用多样化专家提示生成更多答案,扩大正确解空间。
  • 专家间互学机制使性能提升4.89%(Qwen)至11.33%(Llama)。
  • 适合需要强推理的复杂任务,尤其在低奖励场景下表现优。

近期研究表明,基于可验证奖励的强化学习(RLVR)能显著增强大语言模型(LLM)的推理能力。然而,标准RLVR面临奖励稀疏问题:持续错误的答案导致零奖励,无法提供有效学习信号,尤其在复杂任务中。为此,我们提出多专家互学的GRPO框架(MEML-GRPO),利用多样化的专家提示作为系统提示,生成更广泛的响应,大幅提升找到正确解的概率。此外,引入专家间互学机制,促进知识共享与迁移,进一步提升模型在RLVR中的表现。在多个推理基准上的实验表明,MEML-GRPO取得显著改进,使用Qwen时平均性能提升4.89%,使用Llama时达11.33%,有效克服了传统RLVR的核心局限。

原文摘要 · Abstract (English)

Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers provide no learning signal, particularly in challenging tasks. To address this, we propose Multi-Expert Mutual Learning GRPO (MEML-GRPO), an innovative framework that utilizes diverse expert prompts as system prompts to generate a broader range of responses, substantially increasing the likelihood of identifying correct solutions. Additionally, we introduce an inter-expert mutual learning mechanism that facilitates knowledge sharing and transfer among experts, further boosting the model's performance through RLVR. Extensive experiments across multiple reasoning benchmarks show that MEML-GRPO delivers significant improvements, achieving an average performance gain of 4.89% with Qwen and 11.33% with Llama, effectively overcoming the core limitations of traditional RLVR methods.

强化学习推理增强多专家学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。