arXiv:2410.20796cs.CL2024-10被引 7

用多语言低质量数据重写提升大模型预训练效果

Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training

  • 构建优化重写流水线,处理英德意西多语种数据
  • 低质量数据重写后模型性能显著提升,高质量数据增益递减
  • 不同模型家族差异大于模型规模差异,选型需谨慎测试

近期研究表明,将原始数据与合成重写数据结合用于大模型预训练可取得良好效果。本文在C4数据集上复现相关成果,并将其扩展至多语言CulturaX数据集(含英语、德语、意大利语和西班牙语的Oscar子集),通过优化重写流水线,在单语和多语设置下均提升了标准评估基准的表现。我们深入分析了重写流程中基数据集与基础模型的选择,以及模型规模与预训练后性能的关系。研究发现,随着数据质量提高,性能增益逐渐下降;且不同模型家族间的性能差异大于同族不同规模模型之间的差异。这表明选择重写工具前需进行细致评估。此外,我们考察了合成数据对监督微调的影响,结果显示增益存在但不明确,高度依赖评估基准,再次凸显当前评测体系的不足。总体而言,重写多语言低质量数据是扩展大模型预训练数据的有力方向。

原文摘要 · Abstract (English)

Recently published work on rephrasing natural text data for pre-training LLMs has shown promising results when combining the original dataset with the synthetically rephrased data. We build upon previous work by replicating existing results on C4 and extending them with our optimized rephrasing pipeline to the English, German, Italian, and Spanish Oscar subsets of CulturaX. Our pipeline leads to increased performance on standard evaluation benchmarks in both the mono- and multilingual setup. In addition, we provide a detailed study of our pipeline, investigating the choice of the base dataset and LLM for the rephrasing, as well as the relationship between the model size and the performance after pre-training. By exploring data with different perceived quality levels, we show that gains decrease with higher quality. Furthermore, we find the difference in performance between model families to be bigger than between different model sizes. This highlights the necessity for detailed tests before choosing an LLM to rephrase large amounts of data. Moreover, we investigate the effect of pre-training with synthetic data on supervised fine-tuning. Here, we find increasing but inconclusive results that highly depend on the used benchmark. These results (again) highlight the need for better benchmarking setups. In summary, we show that rephrasing multilingual and low-quality data is a very promising direction to extend LLM pre-training data.

大模型预训练数据增强多语言合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。