开源平台OpenDataArena让训练数据价值可衡量、可追溯,推动数据驱动AI的科学化
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
- 构建统一评估流水线,公平比较不同模型与数据集的表现
- 分析超120个数据集,发现数据复杂度与任务性能存在权衡关系
- 适合关注数据质量、模型可复现性及数据工程的研究者
大型语言模型的快速发展依赖于后训练数据的质量与多样性。然而,模型虽有严格评测,其训练数据却仍是黑箱——成分模糊、来源不明、缺乏系统评估,阻碍了可复现性,并掩盖了数据特征与模型行为间的因果关系。为此,我们提出OpenDataArena(ODA),一个全面开放的平台,用于评估后训练数据的内在价值。ODA包含四大支柱:(i) 统一的训练-评估流水线,确保在多种模型(如Llama、Qwen)和领域间实现公平、开放的对比;(ii) 多维度评分框架,在数十个维度上刻画数据质量;(iii) 交互式数据溯源工具,可视化数据谱系并拆解来源组件;(iv) 完全开源的训练、评估与评分工具包,促进数据研究。在涵盖22个基准、超过120个训练数据集的实验中,通过600余次训练运行和4000万条处理数据点,我们揭示出数据复杂度与任务性能之间的非平凡权衡,通过谱系追踪识别出主流基准中的冗余内容,并绘制出数据集间的基因谱系关系。所有结果、工具与配置均公开,旨在推动高质量数据评估的普及。不同于单纯扩充排行榜,ODA致力于从试错式数据构建转向以数据为中心的科学范式,为数据混合规律与基础模型战略构成提供严谨研究基础。
原文摘要 · Abstract (English)
The rapid evolution of Large Language Models (LLMs) is predicated on the quality and diversity of post-training datasets. However, a critical dichotomy persists: while models are rigorously benchmarked, the data fueling them remains a black box--characterized by opaque composition, uncertain provenance, and a lack of systematic evaluation. This opacity hinders reproducibility and obscures the causal link between data characteristics and model behaviors. To bridge this gap, we introduce OpenDataArena (ODA), a holistic and open platform designed to benchmark the intrinsic value of post-training data. ODA establishes a comprehensive ecosystem comprising four key pillars: (i) a unified training-evaluation pipeline that ensures fair, open comparisons across diverse models (e.g., Llama, Qwen) and domains; (ii) a multi-dimensional scoring framework that profiles data quality along tens of distinct axes; (iii) an interactive data lineage explorer to visualize dataset genealogy and dissect component sources; and (iv) a fully open-source toolkit for training, evaluation, and scoring to foster data research. Extensive experiments on ODA--covering over 120 training datasets across multiple domains on 22 benchmarks, validated by more than 600 training runs and 40 million processed data points--reveal non-trivial insights. Our analysis uncovers the inherent trade-offs between data complexity and task performance, identifies redundancy in popular benchmarks through lineage tracing, and maps the genealogical relationships across datasets. We release all results, tools, and configurations to democratize access to high-quality data evaluation. Rather than merely expanding a leaderboard, ODA envisions a shift from trial-and-error data curation to a principled science of Data-Centric AI, paving the way for rigorous studies on data mixing laws and the strategic composition of foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。