将开放任务转化为可验证选择题,让大模型在无标准答案时也能强化推理能力。
Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
- 把开放性问题转为可验证的选择题形式进行强化学习训练。
- 在7个开放任务上平均比传统奖励模型方法提升3.29分。
- 适合需要高质量推理但缺乏明确答案的场景,如创意写作与主观问答。
基于可验证奖励的强化学习(RLVR)在提升大语言模型(LLMs)的推理能力方面展现出巨大潜力。然而,其成功目前主要局限于数学和编程等具有清晰、可自动验证结果的领域。对于开放性任务(如创意写作和主观问答),由于缺乏可验证的解,强化学习仍依赖于奖励模型。这引出一个关键问题:如何在没有明确真实答案的情况下,将RLVR扩展到开放性任务以增强推理?为此,我们提出一种新训练策略——基于可验证多选题重构的强化学习(VMR-RLVR),将开放性数据重构为可验证的多选题格式,从而在无显式真值条件下实现有效训练。多个基准测试的结果验证了该方法的有效性。值得注意的是,在7个开放性任务基准上,我们的VMR-RLVR训练相较使用奖励模型的强化学习平均提升3.29分。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards(RLVR) has demonstrated great potential in enhancing the reasoning capabilities of large language models (LLMs). However, its success has thus far been largely confined to the mathematical and programming domains with clear and automatically checkable outcomes. Reinforcement learning on open-ended tasks (e.g., creative writing and subjective Q&A) continues to rely on reward models due to the absence of verifiable solutions. This raises a key question: how can we extend RLVR to strengthen reasoning in open-ended tasks regardless of the absence of the unambiguous ground truth? To overcome this challenge, we introduce Verifiable Multiple-Choice Reformulation for Reinforcement Learning from Verifiable Rewards (VMR-RLVR), a novel training strategy that restructures open-ended data into verifiable multiple-choice formats, enabling effective training even in the absence of explicit ground truth. Experimental results on multiple benchmarks validate the effectiveness of our method in improving LLM performance on open-ended tasks. Notably, across seven open-ended benchmarks, our VMR-RLVR training delivers an average gain of 3.29 points over the RL with reward model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。