arXiv:2604.20549cs.CLcs.AI2026-04中稿 · the 3rd Workshop o…

用多语言数据提升低资源语言质量分类,让高资源语言帮低资源语言筛选好数据

Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection

  • 利用多语言嵌入空间中质量标记的跨语言一致性,实现高质量数据迁移
  • 在1030亿词训练下,法语准确率提升1.2%,低资源语言表现媲美单语基线
  • 需通过第三四分位采样或保留率调优,才能充分挖掘多语言信号

随着大语言模型规模扩大,数据筛选从追求体量转向优化信噪比。然而,许多语言缺乏足够的高质量数据来训练可靠的分类器。本文研究嵌入空间中质量标记的跨语言一致性,提出高资源语言可为低资源语言提供过滤支持。在1030亿词、10亿参数模型上评估多种策略,包括跨语言迁移、第三四分位采样(Q3)和保留率调优。结果表明,大规模多语言数据池化在排名稳定性和综合准确率上均优于单语基线,法语综合归一化准确率提升1.2%,低资源语言表现达到或超过单语基线。但仅靠规模无法保证稳定性;对法语等高资源语言,需通过Q3采样或调整保留率优化决策边界,以充分发挥多语言信号优势。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) scale, data curation has shifted from maximizing volume to optimizing the signal-to-noise ratio by performing quality filtering. However, for many languages, native high quality data is insufficient to train robust quality classifiers. This work investigates the idea that quality markers in embedding space may show cross-lingual consistency, which would allow high-resource languages to subsidize the filtering of low-resource ones. We evaluate various filtering strategies, including cross-lingual transfer, third quartile sampling (Q3), and retention rate tuning. Our results demonstrate that massive multilingual pooling frequently outperforms monolingual baselines in both rank stability and aggregate accuracy for a 1B model trained on 103B tokens, delivering gains for high resource languages (1.2% increase in aggregate normalized accuracy for French) and matching or exceeding monolingual baselines for low-resource languages. However, we find that scale alone does not guarantee stability. Furthermore, for high-resource languages like French, we show that refining the decision boundary through third quartile sampling (Q3) or tuning the retention rate is necessary to fully leverage the multilingual signal.

多语言数据筛选质量分类迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。