arXiv:2506.04178cs.LG2025-06被引 209

开源推理数据集助力模型性能超越大厂闭源方案

OpenThoughts: Data Recipes for Reasoning Models

  • 构建公开推理数据集,通过可控实验优化生成流程
  • 70亿参数模型在多项基准上达领先水平,最高提升20.5个百分点
  • 适合研究者复现与改进推理模型,推动开放研究

推理模型在数学、代码和科学等任务上进展迅速,但最佳训练方法仍不明确,因顶尖模型多依赖闭源数据。为此,OpenThoughts项目致力于构建开源推理数据集。基于初步探索,OpenThoughts2-1M数据集训练出的OpenThinker2-32B模型,在AIME和LiveCodeBench等标准评测中达到DeepSeek-R1-Distill-32B水平。随后通过1000+组受控实验系统优化数据生成流程,形成OpenThoughts3。将该流程扩展至120万样本,并以QwQ-32B为教师模型,训练出OpenThoughts3-7B,实现新纪录:AIME 2025得53%,LiveCodeBench 06/24-01/25得51%,GPQA Diamond得54%,相比DeepSeek-R1-Distill-Qwen-7B分别提升15.3、17.2、20.5个百分点。所有数据集与模型均开放于https://openthoughts.ai。

原文摘要 · Abstract (English)

Reasoning models have made rapid progress on many benchmarks involving math, code, and science. Yet, there are still many open questions about the best training recipes for reasoning since state-of-the-art models often rely on proprietary datasets with little to no public information available. To address this, the goal of the OpenThoughts project is to create open-source datasets for training reasoning models. After initial explorations, our OpenThoughts2-1M dataset led to OpenThinker2-32B, the first model trained on public reasoning data to match DeepSeek-R1-Distill-32B on standard reasoning benchmarks such as AIME and LiveCodeBench. We then improve our dataset further by systematically investigating each step of our data generation pipeline with 1,000+ controlled experiments, which led to OpenThoughts3. Scaling the pipeline to 1.2M examples and using QwQ-32B as teacher yields our OpenThoughts3-7B model, which achieves state-of-the-art results: 53% on AIME 2025, 51% on LiveCodeBench 06/24-01/25, and 54% on GPQA Diamond - improvements of 15.3, 17.2, and 20.5 percentage points compared to the DeepSeek-R1-Distill-Qwen-7B. All of our datasets and models are available on https://openthoughts.ai.

推理模型开源数据模型训练AIME

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。