构建400亿文本标记的多模态数学预训练数据集,提升小模型数学推理能力。
InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning

- 从CommonCrawl提取2400万网页,构建图文交错的高质量数学数据集
- 13亿参数模型仅用400亿文本标记即达深求数学1.3B模型水平
- 开源数据集推动开源多模态数学模型性能新纪录
大规模高质量数据集的预训练对提升大语言模型(LLMs)的推理能力至关重要,尤其在数学等专业领域。尽管重要性已被认可,当前多模态大模型(MLLMs)领域仍缺乏专门针对数学推理的开源预训练数据集。为此,我们提出InfiMM-WebMath-40B,一个由2400万网页、8500万图像链接和400亿文本标记组成的高质量图文交错数据集,所有数据均从CommonCrawl中精心提取与过滤。我们详细描述了数据收集与处理流程。为验证其有效性,我们在纯文本和多模态场景下进行了评估:在纯文本基准上,尽管仅使用400亿文本标记,我们的13亿参数模型性能已媲美使用1200亿标记的DeepSeekMath-1.3B;而在多模态数学基准如MathVerse和We-Math上,引入该数据集后,我们的模型成为开源模型中的新标杆。数据集已开源:https://huggingface.co/datasets/Infi-MM/InfiMM-WebMath-40B。
原文摘要 · Abstract (English)
Pre-training on large-scale, high-quality datasets is crucial for enhancing the reasoning capabilities of Large Language Models (LLMs), especially in specialized domains such as mathematics. Despite the recognized importance, the Multimodal LLMs (MLLMs) field currently lacks a comprehensive open-source pre-training dataset specifically designed for mathematical reasoning. To address this gap, we introduce InfiMM-WebMath-40B, a high-quality dataset of interleaved image-text documents. It comprises 24 million web pages, 85 million associated image URLs, and 40 billion text tokens, all meticulously extracted and filtered from CommonCrawl. We provide a detailed overview of our data collection and processing pipeline. To demonstrate the robustness of InfiMM-WebMath-40B, we conducted evaluations in both text-only and multimodal settings. Our evaluations on text-only benchmarks show that, despite utilizing only 40 billion tokens, our dataset significantly enhances the performance of our 1.3B model, delivering results comparable to DeepSeekMath-1.3B, which uses 120 billion tokens for the same model size. Nevertheless, with the introduction of our multi-modal math pre-training dataset, our models set a new state-of-the-art among open-source models on multi-modal math benchmarks such as MathVerse and We-Math. We release our data at https://huggingface.co/datasets/Infi-MM/InfiMM-WebMath-40B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。