用英语数据质量信号,自动筛选多语言训练数据。
MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
- 通过英语评分器对比学习,生成统一的跨语言文档质量评分。
- 在17种语言上提升模型表现,知识类任务准确率显著提高。
- 适合需要高质量多语言预训练数据的研究者使用。
数据质量是影响大模型性能的关键因素,但现有基于模型的数据选择方法几乎仅针对英语。本文提出MuRating,一个可扩展的框架,将高质量英语数据质量信号迁移至17种目标语言。MuRating通过成对比较聚合多个英语评分器,学习统一的文档质量得分,并利用翻译将这些判断投射到单语、跨语言及平行文本对上,训练一个多语言评估模型。该方法应用于网络数据,选取英文与多语言内容平衡的子集,用于预训练12亿参数的LLaMA模型。相比强基线(如QuRater、AskLLM、DCLM等),本方法在英文基准和多语言评估中均取得更高平均准确率,尤其在知识密集型任务上增益显著。进一步分析了翻译保真度、选择偏差及叙事类内容欠代表问题,指明未来研究方向。
原文摘要 · Abstract (English)
Data quality is a critical driver of large language model performance, yet existing model-based selection methods focus almost exclusively on English. We introduce MuRating, a scalable framework that transfers high-quality English data-quality signals into a single rater for 17 target languages. MuRating aggregates multiple English "raters" via pairwise comparisons to learn unified document-quality scores,then projects these judgments through translation to train a multilingual evaluator on monolingual, cross-lingual, and parallel text pairs. Applied to web data, MuRating selects balanced subsets of English and multilingual content to pretrain a 1.2 B-parameter LLaMA model. Compared to strong baselines, including QuRater, AskLLM, DCLM and so on, our approach boosts average accuracy on both English benchmarks and multilingual evaluations, with especially large gains on knowledge-intensive tasks. We further analyze translation fidelity, selection biases, and underrepresentation of narrative material, outlining directions for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。