arXiv:2503.23829cs.CL2025-03被引 163

用可验证奖励提升大模型在多领域的推理能力,无需结构化答案。

Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

  • 基于生成评分构建软性奖励信号,突破二元验证局限。
  • 7B小模型即可训练跨领域奖励模型,显著优于主流开源对齐模型。
  • 适用于医学、化学等开放域场景,增强RL在复杂环境中的实用性。

强化学习中可验证奖励(RLVR)在数学推理与编程任务中表现优异,尤其当存在结构化参考答案时。然而其在更广泛、非结构化领域(如医学、化学、心理学、经济学、教育学)的应用尚未探索。本研究发现,在专家撰写参考答案的前提下,不同大模型对宽领域任务的二元验证判断具高度一致性。受此启发,我们提出一种生成式评分方法,生成基于模型的软性奖励信号,以应对自由文本、无结构回答场景下的验证挑战。进一步证明,仅使用相对较小的7B参数量模型,即可训练出跨领域的生成式奖励模型,无需大量领域标注数据。通过全面实验验证,该框架在自由形式设置下显著超越Qwen2.5-72B与DeepSeek-R1-Distill-Qwen-32B等先进开源对齐模型。本方法显著提升了RLVR在复杂、噪声标签环境下的鲁棒性、灵活性与可扩展性,为实际强化学习应用迈出关键一步。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has demonstrated significant success in enhancing mathematical reasoning and coding performance of large language models (LLMs), especially when structured reference answers are accessible for verification. However, its extension to broader, less structured domains remains unexplored. In this work, we investigate the effectiveness and scalability of RLVR across diverse real-world domains including medicine, chemistry, psychology, economics, and education, where structured reference answers are typically unavailable. We reveal that binary verification judgments on broad-domain tasks exhibit high consistency across various LLMs provided expert-written reference answers exist. Motivated by this finding, we utilize a generative scoring technique that yields soft, model-based reward signals to overcome limitations posed by binary verifications, especially in free-form, unstructured answer scenarios. We further demonstrate the feasibility of training cross-domain generative reward models using relatively small (7B) LLMs without the need for extensive domain-specific annotation. Through comprehensive experiments, our RLVR framework establishes clear performance gains, significantly outperforming state-of-the-art open-source aligned models such as Qwen2.5-72B and DeepSeek-R1-Distill-Qwen-32B across domains in free-form settings. Our approach notably enhances the robustness, flexibility, and scalability of RLVR, representing a substantial step towards practical reinforcement learning applications in complex, noisy-label scenarios.

强化学习大模型可验证奖励跨领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。