用闭环反馈优化训练数据,让模型学得更准更省。
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
- 基于价值评估反馈,自动迭代优化数据集构造
- 数学推理数据集在AIME/HMMT上达最新水平
- 适合追求数据效率与模型性能的开发者
大规模语言模型后训练中的监督微调(SFT)数据集构建至关重要却缺乏理论指导,现有方法多依赖经验拼凑,未能系统理解单个样本对模型表现的影响。本文提出一种从随意筛选转向闭环数据工程的新范式——OpenDataArena(ODA),通过价值锚定排名与多维分析,将评估结果转化为指导数据构建的反馈信号。我们基于此框架构建了两个新数据集: extbf{ODA-Math-460k},采用新颖的两阶段难度感知流程,在AIME和HMMT等数学推理基准上达到当前最优表现; extbf{ODA-Mixture (100k & 500k)},通过“锚定-补丁”策略构建的多领域指令数据集,显著优于更大规模的开源基线。实证结果表明,ODA驱动的数据集在特定领域推理与通用能力上均有提升,且具备更高数据效率,验证了以透明评估为引擎的数据中心型AI转型可行性。
原文摘要 · Abstract (English)
The construction of Supervised Fine-Tuning (SFT) datasets is a critical yet under-theorized stage in the post-training of Large Language Models (LLMs), as prevalent practices often rely on heuristic aggregation without a systematic understanding of how individual samples contribute to model performance. In this report, we propose a paradigm shift from ad-hoc curation to a closed-loop dataset engineering framework using OpenDataArena (ODA), which leverages value-anchored rankings and multi-dimensional analysis to transform value benchmarking into feedback signals guiding dataset construction. We instantiate this methodology through two new datasets: \textbf{ODA-Math-460k}, a specialized mathematics reasoning dataset that utilizes a novel two-stage difficulty-aware pipeline to achieve State-of-the-Art (SOTA) results on benchmarks such as AIME and HMMT, and \textbf{ODA-Mixture (100k \& 500k)}, a series of multi-domain instruction datasets built via an ``Anchor-and-Patch'' strategy that outperforms significantly larger open-source baselines. Our empirical results demonstrate that ODA-driven datasets significantly improve both domain-specific reasoning and general utility while achieving superior data efficiency, validating a transition toward data-centric AI where transparent evaluation serves as the primary engine for engineering high-quality training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。