arXiv:2510.01459cs.LGcs.CL2025-10被引 11

根据回答长度动态选数据,提升大模型推理训练效果

LSPO: Length-aware Dynamic Sampling for Policy Optimization in LLM Reasoning

  • 依据平均回答长度动态选择训练样本
  • 在多个模型和数据集上稳定提升推理能力
  • 揭示长度信号对强化学习优化的关键作用

自 Deepseek-R1 发布以来,基于可验证奖励的强化学习(RLVR)已成为训练大语言模型进行推理任务的核心方法。近期研究主要聚焦于修改损失函数以提高 RLVR 的效率与效果。本文受大模型过度思考现象的启发,提出长度感知的策略优化动态采样方法(LSPO),一种新型元 RLVR 算法,能够在每一步根据平均响应长度动态选择训练数据。我们在多个基础模型和数据集上评估了 LSPO,结果表明其能持续提升学习有效性。此外,我们还进行了详尽的消融实验,探究将长度信号融入动态采样的不同方式,进一步揭示潜在研究方向。

原文摘要 · Abstract (English)

Since the release of Deepseek-R1, reinforcement learning with verifiable rewards (RLVR) has become a central approach for training large language models (LLMs) on reasoning tasks. Recent work has largely focused on modifying loss functions to make RLVR more efficient and effective. In this paper, motivated by studies of overthinking in LLMs, we propose Length-aware Sampling for Policy Optimization (LSPO), a novel meta-RLVR algorithm that dynamically selects training data at each step based on the average response length. We evaluate LSPO across multiple base models and datasets, demonstrating that it consistently improves learning effectiveness. In addition, we conduct a detailed ablation study to examine alternative ways of incorporating length signals into dynamic sampling, offering further insights and highlighting promising directions for future research.

强化学习大模型推理动态采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。