arXiv:2504.02807cs.CLcs.AI2025-04被引 43

构建了迄今最大的开源数学预训练数据集,助力大模型数学推理能力提升。

MegaMath: Pushing the Limits of Open Math Corpora

  • 从网页、代码库和合成数据三方面整合高质量数学数据
  • 总计3710亿token,规模与质量均领先现有公开数据集
  • 适合研究数学推理、LLM预训练的学者与开发者使用

数学推理是人类智能的核心,也是评估大语言模型(LLM)高级能力的关键基准。然而,当前研究仍缺乏一个面向数学中心型LLM预训练的大规模、高质量、开源数据集。本文提出MegaMath,通过三种策略构建:(1) 重采网络数据:利用数学导向的HTML优化、fasttext过滤与去重,从Common Crawl重新提取数学文档;(2) 回溯代码数据:从大型代码语料库Stack-V2中识别高质量数学相关代码,提升数据多样性;(3) 探索合成数据:基于网页或代码数据生成问答式文本、数学代码及交错文本-代码块。通过系统消融实验验证有效性,MegaMath最终提供3710亿token数据,在规模与质量上均超越现有公开数学预训练数据集。

原文摘要 · Abstract (English)

Mathematical reasoning is a cornerstone of human intelligence and a key benchmark for advanced capabilities in large language models (LLMs). However, the research community still lacks an open, large-scale, high-quality corpus tailored to the demands of math-centric LLM pre-training. We present MegaMath, an open dataset curated from diverse, math-focused sources through following practices: (1) Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on the Internet. (2) Recalling Math-related code data: We identified high quality math-related code from large code training corpus, Stack-V2, further enhancing data diversity. (3) Exploring Synthetic data: We synthesized QA-style text, math-related code, and interleaved text-code blocks from web data or code data. By integrating these strategies and validating their effectiveness through extensive ablations, MegaMath delivers 371B tokens with the largest quantity and top quality among existing open math pre-training datasets.

数学推理数据集LLM预训练开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。