用十级数据进化框架提升科学数据价值,让模型越用越强。
Data Darwinism Part I: Unlocking the Value of Scientific Data for Pre-training
- 构建十级数据进化体系,让模型生成更高质量训练数据。
- 在20+任务上超越基线2.12~8.40分,领域对齐任务提升显著。
- 开源数据集与模型,支持可复现的模型数据协同进化。
数据质量决定基础模型性能,但系统化处理框架仍缺。本文提出Data Darwinism,一个十级分类体系(L0-L9),描述数据与模型的共同进化:先进模型能生成下一代系统的优质数据。我们在科学文献中验证该框架,构建了900B-token的Darwin-Science语料库(L0-L5)。发现原始科学文本存在可学习性差距,通过使用前沿大模型在L4(生成精炼)和L5(认知补全)层级显式解释推理与术语,有效填补该空白。为确保严格溯源,我们从头预训练daVinci-origin-3B/7B模型,排除科学内容以构建无污染基线。经600B tokens持续预训练后,Darwin-Science在20多个基准测试中分别领先基线+2.12(3B)和+2.95(7B)分,在领域对齐任务上进一步提升至+5.60和+8.40分。系统推进至L5带来+1.36的总增益,证实高级处理能释放数据潜在价值。论文发布Darwin-Science语料库及daVinci-origin模型,支持原则化、共进化的研发。
原文摘要 · Abstract (English)
Data quality determines foundation model performance, yet systematic processing frameworks are lacking. We introduce Data Darwinism, a ten-level taxonomy (L0-L9) that conceptualizes data-model co-evolution: advanced models produce superior data for next-generation systems. We validate this on scientific literature by constructing Darwin-Science, a 900B-token corpus (L0-L5). We identify a learnability gap in raw scientific text, which we bridge via L4 (Generative Refinement) and L5 (Cognitive Completion) using frontier LLMs to explicate reasoning and terminology. To ensure rigorous attribution, we pre-trained daVinci-origin-3B/7B models from scratch, excluding scientific content to create contamination-free baselines. After 600B tokens of continued pre-training, Darwin-Science outperforms baselines by +2.12 (3B) and +2.95 (7B) points across 20+ benchmarks, rising to +5.60 and +8.40 points on domain-aligned tasks. Systematic progression to L5 yields a +1.36 total gain, confirming that higher-level processing unlocks latent data value. We release the Darwin-Science corpus and daVinci-origin models to enable principled, co-evolutionary development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。