arXiv:2603.01907cs.LGcs.CL2026-03被引 3

用信息量筛选数据,让强化学习训练更高效准确

Efficient RLVR Training via Weighted Mutual Information Data Selection

  • 基于加权互信息构建稳定的数据选择分数,避免噪声干扰
  • 在数学与规划任务上提升1.41分,推理能力提升1.01分
  • 适用于多轮奖励的强化学习场景,计算开销极小

强化学习在提升大语言模型推理与对齐能力中起核心作用,但其效率高度依赖训练数据的选择。现有在线选择策略主要依赖难度启发式,偏好中等成功率样本,隐含将难度等同于信息量,忽略了由证据不足引发的认知不确定性。本文提出InSight——一种基于信息引导的数据采样方法,其目标函数为加权互信息。通过贝叶斯建模数据结果的成功率潜变量,我们发现预期不确定性减少可分解为难度和证据依赖两部分,揭示了仅凭难度选择的局限性。InSight采用样本成功概率的均值信念作为采集得分,而非噪声采样结果,并自然扩展至常见于强化学习与验证奖励(RLVR)中的多轮设置。大量实验表明,InSight持续达到顶尖性能,显著提升训练效率:在规划与数学基准上平均提升1.41分,通用推理能力提升1.01分,最快可提速约2.2倍,且计算开销几乎可忽略。

原文摘要 · Abstract (English)

Reinforcement learning (RL) plays a central role in improving the reasoning and alignment of large language models, yet its efficiency critically depends on how training data are selected. Existing online selection strategies predominantly rely on difficulty-based heuristics, favouring datapoints with intermediate success rates, implicitly equating difficulty with informativeness and neglecting epistemic uncertainty arising from limited evidence. We introduce InSight, an INformation-guided data SamplInG metHod for RL Training, grounded in a weighted mutual information objective. By modeling data outcomes with Bayesian latent success rates, we show that expected uncertainty reduction decomposes into complementary difficulty- and evidence-dependent components, revealing a fundamental limitation of difficulty-only selection. Leveraging this observation, InSight constructs a stable acquisition score based on the mean belief of datapoints' success rather than noisy sampled outcomes, and naturally extends to multi-rollout settings common in reinforcement learning with verifiable rewards (RLVR). Extensive experiments demonstrate that InSight consistently achieves state-of-the-art performance and improves training efficiency, including a +1.41 average gain on Planning & Mathmatics benchmarks, +1.01 improvement on general reasoning, and up to ~2.2x acceleration, with negligible additional computational overhead.

强化学习数据选择信息量大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。