arXiv:2602.12414cs.CL2026-02被引 4

用18维属性标注文档,让大模型训练数据更透明可分析

propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale

  • 用小型多语言模型同时评估18个维度的文档质量
  • 40亿参数模型比更大通用模型更贴近商业级标注者
  • 适合需要精细数据筛选与分析的研究者和工程师

自FineWeb-Edu以来,大模型预训练的数据清洗主要依赖小型分类器生成的单一质量分数。这种单一分数混淆了多个质量维度,无法灵活过滤且缺乏可解释性。我们提出propella-1,一个包含0.6B、1.7B、4B参数的小型多语言LLM家族,可对文本文档进行18项属性标注,涵盖六大类别:核心内容、分类、质量与价值、受众与目的、安全合规、地理相关性。模型支持57种语言,输出符合预定义模式的结构化JSON。在与前沿商业大模型对比中,4B版本的标注一致性高于更大通用模型。我们发布了propella-annotations数据集,包含超过三亿条文档标注,覆盖FineWeb-2、FinePDFs、HPLT 3.0、Nemotron-CC等主流预训练语料。利用这些标注,我们对常用预训练数据集进行了多维度组合分析,揭示出单分值方法无法捕捉的质量、推理深度和内容构成差异。所有模型权重与标注数据均以宽松商用许可发布。

原文摘要 · Abstract (English)

Since FineWeb-Edu, data curation for LLM pretraining has predominantly relied on single scalar quality scores produced by small classifiers. A single score conflates multiple quality dimensions, prevents flexible filtering, and offers no interpretability. We introduce propella-1, a family of small multilingual LLMs (0.6B, 1.7B, 4B parameters) that annotate text documents across 18 properties organized into six categories: core content, classification, quality and value, audience and purpose, safety and compliance, and geographic relevance. The models support 57 languages and produce structured JSON annotations conforming to a predefined schema. Evaluated against a frontier commercial LLM as a reference annotator, the 4B model achieves higher agreement than much larger general-purpose models. We release propella-annotations, a dataset of over three billion document annotations covering major pretraining corpora including data from FineWeb-2, FinePDFs, HPLT 3.0, and Nemotron-CC. Using these annotations, we present a multi-dimensional compositional analysis of widely used pretraining datasets, revealing substantial differences in quality, reasoning depth, and content composition that single-score approaches cannot capture. All model weights and annotations are released under permissive, commercial-use licenses.

数据清洗多属性标注大模型训练可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。