用合成重写提升葡萄牙语预训练数据质量,效果随模型规模增大而增强
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
- 用7B指令微调模型对高质量和低质量数据进行四种风格重写
- 7B模型在高质量数据上重写后提升3.4分NPM,低质量仅提升0.5分
- 重写主要放大数据质量优势,适合大模型在葡萄牙语任务中使用
通过文档重写生成合成数据已成为提升语言模型预训练的有效方法,但多数研究集中于英语,且未系统控制源数据质量。本文针对葡萄牙语持续预训练开展受控实验,基于带有STEM与教育质量评分的ClassiCC-PT语料库,构建两个10B token的高低质量子集,并分别用7B指令微调模型重写为四种风格,每条件生成约40B合成数据。在两个英文基线模型(1.1B和7B参数)上训练并评估在PoETa V2(44项任务的综合葡萄牙语基准)上的表现。7B模型下,重写高质量数据相比原始数据提升3.4 NPM,而重写低质量数据仅提升0.5 NPM;1.1B模型下该效应较弱,未修改的低质量数据与重写的高质量数据表现相近。结果表明,合成重写主要作为质量倍增器而非数据筛选替代品,且其效果具有规模依赖性。
原文摘要 · Abstract (English)
Synthetic data generation through document rewriting has emerged as a promising technique for improving language model pretraining, yet most studies focus on English and do not systematically control for the quality of the source data being rewritten. We present a controlled study of how synthetic rewriting interacts with source data quality in the context of Portuguese continued pretraining. Starting from ClassiCC-PT, a Portuguese corpus annotated with STEM and Educational quality scores, we construct two 10B-token subsets at different quality levels and rewrite each into four styles using a 7B instruction-tuned model, producing approximately 40B tokens of synthetic data per condition. We train two English-centric base models (1.1B and 7B parameters) on each condition and evaluate on PoETa V2, a comprehensive 44-task Portuguese benchmark. At the 7B scale, rewriting high-quality data yields a +3.4 NPM gain over the same data unmodified, while rewriting low-quality data provides only +0.5 NPM. At the 1.1B scale, this interaction is weaker, with unmodified low-quality data performing comparably to rewritten high-quality data. Our results demonstrate that synthetic rewriting acts primarily as a quality multiplier rather than a substitute for data curation, and that this effect is scale-dependent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。