arXiv:2505.23433cs.LG2025-05NeurIPS被引 39

提升大模型推理多样性,让答案更丰富可靠

Diversity-Aware Policy Optimization for Large Language Model Reasoning

  • 在强化学习中引入分层多样性目标,只优化正样本
  • 4个数学推理基准平均提效3.5%,解法更多样
  • 适合追求高质量、多角度推理的开发者和研究者

大语言模型(LLM)的推理能力快速提升,尤其在DeepSeek R1发布后,数据质量和强化学习(RL)算法成为研究热点。尽管多样性在强化学习中至关重要,其对LLM推理的影响仍被忽视。本文系统研究了基于强化学习训练中多样性的作用,提出一种新型多样性感知策略优化方法。在12个大模型上评估发现,高性能模型的解法多样性与新提出的「Potential at k」指标(衡量推理潜力)呈强正相关。据此,我们设计了基于词元级别的多样性机制,并将其转化为可执行目标,仅应用于正样本。集成至R1-zero训练框架后,该方法在4个数学推理基准上实现平均3.5%的性能提升,同时生成更多样、更鲁棒的解法。

原文摘要 · Abstract (English)

The reasoning capabilities of large language models (LLMs) have advanced rapidly, particularly following the release of DeepSeek R1, which has inspired a surge of research into data quality and reinforcement learning (RL) algorithms. Despite the pivotal role diversity plays in RL, its influence on LLM reasoning remains largely underexplored. To bridge this gap, this work presents a systematic investigation into the impact of diversity in RL-based training for LLM reasoning, and proposes a novel diversity-aware policy optimization method. Across evaluations on 12 LLMs, we observe a strong positive correlation between the solution diversity and Potential at k (a novel metric quantifying an LLM's reasoning potential) in high-performing models. This finding motivates our method to explicitly promote diversity during RL training. Specifically, we design a token-level diversity and reformulate it into a practical objective, then we selectively apply it to positive samples. Integrated into the R1-zero training framework, our method achieves a 3.5 percent average improvement across four mathematical reasoning benchmarks, while generating more diverse and robust solutions.

大模型推理强化学习多样性优化数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。