arXiv:2510.07169cs.CL2025-10EMNLP被引 5

优化数学推理能力,好数据比多数据更重要

More Data or Better Data? A Critical Analysis of Data Selection and Synthesis for Mathematical Reasoning

  • 用可解释格式和强模型蒸馏提升数据质量
  • 高质量数据在推理任务上显著优于单纯增加数据量
  • 适合工业场景的数据筛选与合成方法

大型语言模型(LLMs)的推理能力对下游任务至关重要,但高度依赖训练数据质量。尽管已有多种数据构建方法,其在真实应用中的有效性仍缺乏深入研究。本文对开源数学推理数据集及数据合成技术进行了全面分析,采用统一训练与部署流程进行评估。结果表明,将数据结构化为更易理解的格式,或通过更强模型进行知识蒸馏,往往比单纯扩大数据规模更为有效。本研究提炼出实用的数据选择策略,为工业级应用提供可落地的方法,助力低成本数据治理与模型能力持续提升。期望推动对‘更多数据’与‘更好数据’平衡机制的进一步探索。

原文摘要 · Abstract (English)

The reasoning capabilities of Large Language Models (LLMs) play a critical role in many downstream tasks, yet depend strongly on the quality of training data. Despite various proposed data construction methods, their practical utility in real-world pipelines remains underexplored. In this work, we conduct a comprehensive analysis of open-source datasets and data synthesis techniques for mathematical reasoning, evaluating them under a unified pipeline designed to mirror training and deployment scenarios. We further distill effective data selection strategies and identify practical methods suitable for industrial applications. Our findings highlight that structuring data in more interpretable formats, or distilling from stronger models often outweighs simply scaling up data volume. This study provides actionable guidance for integrating training data to enhance LLM capabilities, supporting both cost-effective data curation and scalable model enhancement. We hope this work will inspire further research on how to balance "more data" versus "better data" for real-world reasoning tasks.

数学推理数据质量模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。