arXiv:2604.07747cs.AIcs.CL2026-04

通过智能提示合成与渐进去提示,提升数学强化学习的解题泛化能力。

Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing

  • 用学生风格生成匹配的验证提示,缓解师生分布差异。
  • 分难度梯度降低提示暴露,保持无提示更新能力。
  • 在大采样数下显著提升性能,适合复杂数学推理训练。

基于可验证奖励的强化学习(RLVR)能在提升低采样数推理准确率的同时缩小解法覆盖范围,且单次通过率(pass@1)的提升不必然带来高采样数(large-k)性能的改善。现有提示方法虽使难题可训练,但未充分解决教师-学生分布不匹配及提示暴露需与无提示评估对齐的问题。本文提出两个组件:分布对齐提示合成(DAHS)根据学生式回答生成验证过的教师提示;反向提示退火(BHA)按难度分桶逐步减少提示暴露,并使用每题提示丢弃以保留无提示更新。在DAPO框架下,使用Qwen3-1.7B-Base和Llama-3.2-1B-Instruct模型,在AIME24、AIME25、AIME26三个基准上评估。Qwen3-1.7B-Base上,本方法在pass@1和pass@2048上均优于DAPO;Llama-3.2-1B-Instruct上,增益集中在大-$k$场景。结果表明,在数学RLVR中,早期通过提示提供可学更新并后期逐步移除提示,是有效的提示支架策略。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) can improve low-$k$ reasoning accuracy while narrowing solution coverage on challenging math questions, and pass@1 gains do not necessarily translate into better large-$k$ performance. Existing hint-based approaches can make challenging questions trainable, but they leave two issues underexplored: teacher-student distribution mismatch and the need to reduce hint exposure to match no-hint evaluation. We address these issues through two components. Distribution-Aligned Hint Synthesis (DAHS) constructs verified teacher hints conditioned on student-style responses. Backward Hint Annealing (BHA) anneals hint exposure across difficulty buckets and uses per-question hint dropout to preserve no-hint updates throughout RL training. We evaluate the method in math RLVR under the DAPO training framework across AIME24, AIME25, and AIME26 using $\texttt{Qwen3-1.7B-Base}$ and $\texttt{Llama-3.2-1B-Instruct}$. On $\texttt{Qwen3-1.7B-Base}$, our method improves both pass@1 and pass@2048 relative to DAPO across the three AIME benchmarks. On $\texttt{Llama-3.2-1B-Instruct}$, the gains are concentrated in the large-$k$ regime. These results suggest that, in math RLVR, hint scaffolding is effective when it restores learnable updates on challenging questions early in training and is then gradually removed before no-hint evaluation.

数学推理强化学习提示工程训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。