arXiv:2510.09001cs.CL2025-10被引 4

动态调整难度权重,让大模型数学推理训练更高效

DARO: Difficulty-Aware Reweighting Policy Optimization

  • 根据模型学习状态实时调整不同难度样本的损失权重
  • 在6个数学基准上超越4个主流基线,收敛更快、效果更好
  • 适合追求高效推理能力训练的大模型研究者

大语言模型的推理能力可通过可验证奖励强化学习(RLVR)显著提升。当前主流方法组相对策略优化(GRPO)虽有效,但其依赖静态或过于简单的样本难度加权机制,无法随模型能力演化自适应调整,导致损失尺度失衡,部分难度水平被过度关注而影响整体性能。为此,我们提出困难感知重加权策略优化(DARO),动态调节各难度组的损失贡献。在Qwen2.5-Math-1.5B、Qwen2.5-Math-7B和Llama3.1-8B上的大量实验表明,DARO在六个数学基准上均优于四个领先基线,实现更快收敛与更优最终表现。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have shown that reasoning ability can be significantly enhanced through Reinforcement Learning with Verifiable Rewards (RLVR). Group Relative Policy Optimization (GRPO) has emerged as the de facto approach for RLVR, inspiring numerous variants. However, our mathematical analysis reveals that these methods are fundamentally weighted variations of GRPO. We provide a unified view, demonstrating that their reliance on static or overly simplistic weighting schemes tied to sample difficulty prevents adaptation to a model's evolving capabilities. This creates a significant loss scale issue, where training disproportionately focuses on certain difficulty levels at the expense of others, hindering overall performance. To address these limitations, we introduce \textbf{Difficulty-Aware Reweighting Policy Optimization (DARO)}, a method that dynamically adjusts the loss contribution of each difficulty group based on the model's learning state. Extensive experiments on Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Llama3.1-8B show that DARO outperforms four leading baselines across six math benchmarks, achieving significantly faster convergence and superior final performance.

强化学习大模型数学推理动态权重

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。