构建超25万道可验证答案的数学题数据集,专为强化学习训练语言模型设计。
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
- 从开源数据中筛选并清洗25万+高质数学题,确保解法唯一且可验证
- 新增4.7万道经系统重构的开放性题目,适配强化学习任务需求
- 比GSM8k和MATH大一个数量级,兼顾质量与多样性,适合不同能力模型
近年来,推理模型的发展使数学成为算法与方法改进的重要测试场景。然而现有开源数学数据集或高质量问题少,或大规模机器生成问题质量不可靠,研究者难以兼顾质量与数量。本文提出Big-Math,包含超过25万道具有可验证答案的高质量数学题,专为强化学习(RL)设计。通过严格筛选、清洗与人工验证,我们从公开数据集中提取满足三大标准的问题:(1)解法唯一可验证;(2)开放性问题;(3)存在闭式解。在此基础上,我们引入4.7万道新题——Big-Math-Reformulated,即通过系统化重构算法将选择题转化为开放题。相比常用开源数据集GSM8k和MATH,Big-Math规模大一个数量级,同时保持高适配性。分析显示,该数据集覆盖广泛问题领域,难度分布均衡,支持各类模型的多样化下游应用。通过平衡数据量与质量,Big-Math为提升大语言模型的推理能力奠定坚实基础。
原文摘要 · Abstract (English)
Increasing interest in reasoning models has led math to become a prominent testing ground for algorithmic and methodological improvements. However, existing open math datasets either contain a small collection of high-quality, human-written problems or a large corpus of machine-generated problems of uncertain quality, forcing researchers to choose between quality and quantity. In this work, we present Big-Math, a dataset of over 250,000 high-quality math questions with verifiable answers, purposefully made for reinforcement learning (RL). To create Big-Math, we rigorously filter, clean, and curate openly available datasets, extracting questions that satisfy our three desiderata: (1) problems with uniquely verifiable solutions, (2) problems that are open-ended, (3) and problems with a closed-form solution. To ensure the quality of Big-Math, we manually verify each step in our filtering process. Based on the findings from our filtering process, we introduce 47,000 new questions with verified answers, Big-Math-Reformulated: closed-ended questions (i.e. multiple choice questions) that have been reformulated as open-ended questions through a systematic reformulation algorithm. Compared to the most commonly used existing open-source datasets for math reasoning, GSM8k and MATH, Big-Math is an order of magnitude larger, while our rigorous filtering ensures that we maintain the questions most suitable for RL. We also provide a rigorous analysis of the dataset, finding that Big-Math contains a high degree of diversity across problem domains, and incorporates a wide range of problem difficulties, enabling a wide range of downstream uses for models of varying capabilities and training requirements. By bridging the gap between data quality and quantity, Big-Math establish a robust foundation for advancing reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。