arXiv:2504.11456cs.CLcs.AI2025-04被引 240

打造高质量数学推理数据集,助力大模型突破复杂问题求解

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

论文配图:DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
图 1 · 摘自论文原文
  • 构建高难度数学题库,覆盖5-9级挑战题,确保题目真实且无重复
  • 经多轮去污染验证,答案可追溯,支持强化学习奖励机制
  • 模型训练后不仅数学能力顶尖,还能迁移到生物物理等跨学科领域

强化学习结合大语言模型在复杂推理中展现潜力,但受限于缺乏大规模、高难度、无污染且可验证的训练数据。为此,我们推出DeepMath-103K,一个涵盖广泛数学主题的大规模数据集,主要包含5-9级高难度题目,经过多基准测试严格去污染,并提供可验证的答案以支持基于规则的强化学习奖励机制。该数据集包含三种不同R1解法,适配监督微调(SFT)等多种训练范式。在该数据集上训练的模型,在多个高难度数学基准上达到领先水平,并展现出向生物学、物理学和化学等领域的泛化能力,证明其广泛有效性。数据已公开:https://huggingface.co/datasets/zwhe99/DeepMath-103K。

原文摘要 · Abstract (English)

Reinforcement learning (RL) with large language models shows promise in complex reasoning. However, its progress is hindered by the lack of large-scale training data that is sufficiently challenging, contamination-free and verifiable. To this end, we introduce DeepMath-103K, a large-scale mathematical dataset designed with high difficulty (primarily levels 5-9), rigorous decontamination against numerous benchmarks, and verifiable answers for rule-based RL reward. It further includes three distinct R1 solutions adaptable for diverse training paradigms such as supervised fine-tuning (SFT). Spanning a wide range of mathematical topics, DeepMath-103K fosters the development of generalizable and advancing reasoning. Notably, models trained on DeepMath-103K achieve state-of-the-art results on challenging mathematical benchmarks and demonstrate generalization beyond math such as biology, physics and chemistry, underscoring its broad efficacy. Data: https://huggingface.co/datasets/zwhe99/DeepMath-103K.

数学推理大模型训练数据集强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。