arXiv:2601.14700cs.CL2026-01ACL被引 1

让大模型生成更多样化答案,同时保持正确性。

DARL: Encouraging Diverse Answers for General Reasoning without Verifiers

  • 用强化学习鼓励模型在参考答案附近生成多样答案。
  • 在13个基准上表现更好,通用任务提升9.5分。
  • 无需额外验证器,适合开放领域写作任务。

基于可验证奖励的强化学习(RLVR)在提升大语言模型推理能力方面表现优异,但其依赖领域特定验证器,限制了在开放通用领域的应用。近期方法如RLPR虽拓展至通用领域,可在更广数据集上训练并优于RLVR,但仍存在过度拟合参考答案的问题,导致输出多样性不足,尤其在开放任务如写作中更为明显。为此,我们提出DARL,一种简单有效的强化学习框架,能在控制偏离范围的前提下鼓励生成多样化答案,同时保持与参考答案的一致性。该框架完全兼容现有通用强化学习方法,无需额外验证器即可无缝集成。在十三个基准上的实验表明,DARL持续提升推理性能。尤为显著的是,相比RLPR,在六个推理基准上平均提升1.3分,在七个通用基准上提升9.5分,充分证明其在提升推理准确率与输出多样性方面的有效性。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated promising gains in enhancing the reasoning capabilities of large language models. However, its dependence on domain-specific verifiers significantly restricts its applicability to open and general domains. Recent efforts such as RLPR have extended RLVR to general domains, enabling training on broader datasets and achieving improvements over RLVR. However, a notable limitation of these methods is their tendency to overfit to reference answers, which constrains the model's ability to generate diverse outputs. This limitation is particularly pronounced in open-ended tasks such as writing, where multiple plausible answers exist. To address this, we propose DARL, a simple yet effective reinforcement learning framework that encourages the generation of diverse answers within a controlled deviation range from the reference while preserving alignment with it. Our framework is fully compatible with existing general reinforcement learning methods and can be seamlessly integrated without additional verifiers. Extensive experiments on thirteen benchmarks demonstrate consistent improvements in reasoning performance. Notably, DARL surpasses RLPR, achieving average gains of 1.3 points on six reasoning benchmarks and 9.5 points on seven general benchmarks, highlighting its effectiveness in improving both reasoning accuracy and output diversity.

强化学习大模型推理多样性生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。