arXiv:2601.04954cs.LGcs.AI2026-01

高精度奖励比多样约束更有效,能显著提升指令跟随能力。

Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following

  • 用纯硬性约束训练,比混合软硬约束更有效。
  • 新方法在5个基准上性能提升13.4%,训练时间减少58%。
  • 适合追求高效对齐与强泛化的模型开发者。

指导遵循(IF)任务中,强化学习通常依赖可验证奖励的规模扩展,普遍认为需混合使用可验证硬约束和不可验证软约束以实现泛化。本文通过系统实证研究挑战这一共识。出人意料的是,仅使用硬约束训练的模型始终优于混合数据集训练的模型。大量实验表明,奖励精度而非约束多样性才是对齐效果的关键驱动因素。语言模型裁判在检测错误响应时召回率低,导致严重奖励劫持,削弱了多样性的价值。注意力机制分析显示,高精度奖励可发展出可迁移的指令遵循元技能。基于此,我们提出一种以奖励精度为核心的简单数据优化策略。在五个基准测试中,该方法性能优于竞争基线13.4%,训练时间减少58%,且保持对指令跟随之外任务的强泛化能力。研究倡导范式转变:从盲目追求数据多样性转向聚焦高精度奖励。

原文摘要 · Abstract (English)

A central belief in scaling reinforcement learning with verifiable rewards for instruction following (IF) tasks is that, a diverse mixture of verifiable hard and unverifiable soft constraints is essential for generalizing to unseen instructions. In this work, we challenge this prevailing consensus through a systematic empirical investigation. Counter-intuitively, we find that models trained on hard-only constraints consistently outperform those trained on mixed datasets. Extensive experiments reveal that reward precision, rather than constraint diversity, is the primary driver of effective alignment. The LLM judge suffers from a low recall rate in detecting false response, which leads to severe reward hacking, thereby undermining the benefits of diversity. Furthermore, analysis of the attention mechanism reveals that high-precision rewards develop a transferable meta-skill for IF. Motivated by these insights, we propose a simple yet effective data-centric refinement strategy that prioritizes reward precision. Evaluated on five benchmarks, our approach outperforms competitive baselines by 13.4\% in performance while achieving a 58\% reduction in training time, maintaining strong generalization beyond instruction following. Our findings advocate for a paradigm shift: moving away from the indiscriminate pursuit of data diversity toward high-precision rewards.

指令跟随强化学习奖励设计模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。