新数学基准挑战大模型推理极限,推动强化学习探索新解法。
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
- 设计新基准MATH-B,筛选难解题型对抗高采样模型
- 80亿参数以下模型在pass@1024下仍表现不佳
- 适合追求深层推理与探索性强化学习的研究者
随着DeepSeek-R1的出现,强化学习(RL)方法在提升数学推理能力方面崭露头角。然而,开源生态显示:通过足够多采样(如pass@1024),现有基础模型已几乎能解决MATH-500和AIME 2024等主流数学基准的所有题目。这表明当前多数RL微调方法仅是优化已有解题模式,而非发现全新解法。而这种“精炼”与强化学习本应激发探索、习得新技能的初衷相悖。为此,我们提出MATH-Beyond(MATH-B)基准,专为突破8B参数以下开源模型在大规模采样下的表现而设计。其题目源自DAPO-Math-17K和DeepScaleR数据集的子集,保持与高中数学主题一致。验证结果显示,如Nemotron-Research-Reasoning-Qwen-1.5B和DeepScaleR-1.5B-Preview等经过RL微调的模型,在pass@1024下表现依然有限,证明现有方法难以应对更难例题。我们希望该基准能推动以探索为导向的强化学习方法,激发模型更深层面的推理能力。数据集已公开于https://huggingface.co/datasets/brendel-group/MATH-Beyond。
原文摘要 · Abstract (English)
With the advent of DeepSeek-R1, a new wave of reinforcement learning (RL) methods has emerged that seem to unlock stronger mathematical reasoning. However, a closer look at the open-source ecosystem reveals a critical limitation: with sufficiently many draws (e.g., $\texttt{pass@1024}$), many existing base models already solve nearly all questions on widely used math benchmarks such as MATH-500 and AIME 2024. This suggests that the RL fine-tuning methods prevalent in the LLM reasoning literature largely sharpen existing solution modes rather than discovering entirely new ones. Such sharpening stands in contrast to the broader promise of RL: to foster exploration and to acquire new skills. To move beyond this plateau, we introduce MATH-Beyond (MATH-B), a benchmark deliberately constructed to defeat common open-source models of up to 8B parameters even under large sampling budgets. Improving performance on our benchmark via RL requires methods that learn to reason in ways that go beyond base model capabilities in repeated sampling. Since the problems are drawn from subsets of DAPO-Math-17K and DeepScaleR datasets, they remain topically equivalent to standard high-school math. Validating our premise, RL fine-tuned models such as Nemotron-Research-Reasoning-Qwen-1.5B and DeepScaleR-1.5B-Preview perform poorly on MATH-B at $\texttt{pass@1024}$, showing how existing approaches fall short on tackling harder instances. We hope MATH-B will catalyze exploration-driven RL approaches that elicit deeper reasoning capabilities. We release MATH-B at https://huggingface.co/datasets/brendel-group/MATH-Beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。