用合规数据训练大模型,性能媲美主流开源项目。
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- 采用许可优先策略,整合公有领域与宽松授权文本
- 1.7B模型用3000亿词元训练,代码数学任务表现超基准线
- 提供分级溯源信息,适合关注法律风险的研究者
我们提出MixtureVitae,一个开放获取的预训练语料库,旨在最小化法律风险的同时实现优异的下游性能。该数据集遵循许可优先、风险可控的采集策略,融合公有领域及宽松授权文本(如CC-BY/Apache),并谨慎引入低风险补充来源(如政府文件和欧盟允许文本挖掘的资源)。采用单阶段预训练方案,高比例集成合成指令与推理数据——这类信号通常在后训练阶段才出现,且在宽松许可网络语料中稀缺。所有数据源按风险等级分为三级,并提供片级来源元数据,支持风险敏感使用。在标准基准测试中,基于开放科学参考训练协议(固定架构与超参数;500亿与3000亿词元预算,模型规模1.3亿至17亿参数),使用MixtureVitae训练的模型持续优于其他许可宽松数据集。在17亿参数/3000亿词元设置下,其性能接近甚至超越FineWeb-Edu与DCLM晚期水平。尤其在MMLU、数学与代码任务上表现突出:17亿参数模型仅用3000亿词元即在GSM8K、HumanEval与MBPP上达到或超过强指令微调基线,而后者使用约11万亿词元。经彻底去污染分析验证,证明以许可优先、高指令与推理密度、分级风险标注的数据,可为大模型训练提供可行且低风险的基石,减少对广泛网络爬取的依赖而不牺牲竞争力。
原文摘要 · Abstract (English)
We present MixtureVitae, an open-access pretraining corpus built to minimize legal risk while providing strong downstream performance. MixtureVitae follows a permissive-first, risk-mitigated sourcing strategy that combines public-domain and permissively licensed text (e.g., CC-BY/Apache) with carefully justified low-risk additions (e.g., government works and EU TDM-eligible sources). MixtureVitae adopts a simple, single-stage pretraining recipe that integrates a large proportion of permissive synthetic instruction and reasoning data-signals typically introduced during post-training and generally scarce in permissive web corpora. We categorize all sources into a three-tier scheme that reflects varying risk levels and provide shard-level provenance metadata to enable risk-aware usage. In controlled experiments using the open-sci-ref training protocol (fixed architectures and hyperparameters; 50B and 300B token budgets across 130M-1.7B parameters), models trained on MixtureVitae consistently outperform other permissive datasets across a suite of standard benchmarks, and at the 1.7B-parameters/300B-tokens setting, they surpass FineWeb-Edu and approach DCLM late in training. Performance is particularly strong on MMLU and on math and code benchmarks: a 1.7B model pretrained on 300B MixtureVitae tokens matches or exceeds a strong 1.7B instruction-tuned baseline on GSM8K, HumanEval, and MBPP, despite using over 36 times fewer tokens (300B vs. ~11T). Supported by a thorough decontamination analysis, these results show that permissive-first data with high instruction and reasoning density, tiered by licensing and provenance-related risk, can provide a practical and risk-mitigated foundation for training capable LLMs, reducing reliance on broad web scrapes without sacrificing competitiveness. Code: https://github.com/ontocord/mixturevitae
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。