24万亿token的结构化网页数据集,助力高效训练高质量语言模型。
Essential-Web v1.0: 24T tokens of organized web data
- 用12类标签对网页文档分类,实现内容精准组织。
- 仅用SQL过滤即在数学、代码等任务上达SOTA水平。
- 适合需要高质量预训练数据的研究者与开发者使用。
数据在语言模型获取技能与知识中起关键作用。现有大规模、结构化预训练数据集的缺失导致数据管道成本高且难获取。本文提出 Essential-Web v1.0,一个包含24万亿标记的网页数据集,每篇文档均标注了涵盖主题、格式、内容复杂度和质量的12类分类标签。标签由经微调的0.5b参数模型EAI-Distill-0.5b生成,其标注一致性与Qwen2.5-32B-Instruct相差仅3%。仅通过SQL风格筛选,即可在数学任务(相对落后SOTA 8.0%)、代码生成(+14.3%)、STEM(+24.5%)和医学(+8.6%)上获得具有竞争力的结果。Essential-Web v1.0 已在 HuggingFace 上公开:https://huggingface.co/datasets/EssentialAI/essential-web-v1.0
原文摘要 · Abstract (English)
Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible data pipelines. We present Essential-Web v1.0, a 24-trillion-token dataset in which every document is annotated with a twelve-category taxonomy covering topic, format, content complexity, and quality. Taxonomy labels are produced by EAI-Distill-0.5b, a fine-tuned 0.5b-parameter model that achieves an annotator agreement within 3% of Qwen2.5-32B-Instruct. With nothing more than SQL-style filters, we obtain competitive web-curated datasets in math (-8.0% relative to SOTA), web code (+14.3%), STEM (+24.5%) and medical (+8.6%). Essential-Web v1.0 is available on HuggingFace: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。