arXiv:2506.12966cs.CL2025-06EMNLP被引 1

数据质量比数量更重要,提升双语模型性能的关键是筛选高质量数据。

Assessing the Role of Data Quality in Training Bilingual Language Models

  • 通过对比双语与单语模型,发现数据质量不均是性能波动主因。
  • 仅用高质量英语数据过滤后,法德中文单语性能提升2-4%。
  • 适合关注多语言模型平衡性与数据清洗的研究者与工程师。

双语和多语言语言模型为跨语言、跨用户扩展自然语言处理系统提供了前景。然而,现有研究显示,增加语言数量可能使某些语言(如英语)性能下降,而另一些数据稀缺语言的性能则提升。本文通过对比双语与单语模型,探究其不一致性的成因,发现数据质量不均是导致双语设置下性能下降的主要因素,而非数据量。为此,提出一种简单有效的数据过滤策略:仅使用高质量英语数据进行双语训练。该方法在法语、德语和中文上应用后,单语性能提升2-4%,双语模型性能差距缩小至1%。结果凸显了多语言预训练中数据质量被忽视的重要性,并提供了一种实用的性能平衡方案。

原文摘要 · Abstract (English)

Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly between languages as prior works show that adding more languages can degrade performance for some languages (such as English), while improving others (typically more data constrained languages). In this work, we investigate causes of these inconsistencies by comparing bilingual and monolingual language models. Our analysis reveals that unequal data quality, not just data quantity, is a major driver of performance degradation in bilingual settings. We propose a simple yet effective data filtering strategy to select higher-quality bilingual training data with only high quality English data. Applied to French, German, and Chinese, our approach improves monolingual performance by 2-4% and reduces bilingual model performance gaps to 1%. These results highlight the overlooked importance of data quality in multilingual pretraining and offer a practical recipe for balancing performance.

多语言模型数据质量语言对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。