arXiv:2505.14147cs.AI2025-05被引 3

用可验证奖励生成高质量推理题,提升大模型解题能力

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning

  • 设计自对齐原则,确保题目难度与逻辑严谨性
  • 在GPQA上训练后,复杂推理准确率显著提升
  • 适合需要强逻辑推理能力的模型训练场景

在STEM领域用强化学习训练大推理模型(LRMs)受限于高质量、多样化且可验证题目的稀缺。现有合成方法如思维链提示常生成过于简单或不可验证的数据,限制模型在复杂任务上的进步。为此,我们提出SHARP,一种统一的高质对齐推理题合成方法,支持可验证奖励的强化学习(RLVR)。SHARP包含一组自对齐原则——针对研究生及奥赛级别难度、严格逻辑一致性、答案明确可验证——以及结构化的三阶段框架(对齐、实例化、推理),保障主题多样性和细粒度控制。通过先进大模型推断并验证难题,再利用强化学习循环优化模型推理能力。在GPQA等基准测试中,经SHARP增强训练的模型显著优于现有方法,复杂推理准确率大幅提升,使模型性能更接近专家水平。贡献包括SHARP策略、框架设计、端到端实现及有效性实验评估。

原文摘要 · Abstract (English)

Training large reasoning models (LRMs) with reinforcement learning in STEM domains is hindered by the scarcity of high-quality, diverse, and verifiable problem sets. Existing synthesis methods, such as Chain-of-Thought prompting, often generate oversimplified or uncheckable data, limiting model advancement on complex tasks. To address these challenges, we introduce SHARP, a unified approach to Synthesizing High-quality Aligned Reasoning Problems for LRMs reinforcement learning with verifiable rewards (RLVR). SHARP encompasses a strategic set of self-alignment principles -- targeting graduate and Olympiad-level difficulty, rigorous logical consistency, and unambiguous, verifiable answers -- and a structured three-phase framework (Alignment, Instantiation, Inference) that ensures thematic diversity and fine-grained control over problem generation. We implement SHARP by leveraging a state-of-the-art LRM to infer and verify challenging STEM questions, then employ a reinforcement learning loop to refine the model's reasoning through verifiable reward signals. Experiments on benchmarks such as GPQA demonstrate that SHARP-augmented training substantially outperforms existing methods, markedly improving complex reasoning accuracy and pushing LRM performance closer to expert-level proficiency. Our contributions include the SHARP strategy, framework design, end-to-end implementation, and experimental evaluation of its effectiveness in elevating LRM reasoning capabilities.

推理增强强化学习题库生成STEM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。