arXiv:2504.04152cs.CL2025-04被引 11

对比36种多语言持续预训练配置,发现代码数据能提升低资源语言表现

Rethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources

  • 用双语和代码数据混合进行持续预训练,提升多语言分类能力
  • 加入编程代码使低资源语言分类准确率提升,但生成质量略有下降
  • 不同语言类型对跨语言迁移影响复杂,需系统研究才能优化策略

大型语言模型在不同语言间性能差异显著,主要惠及高资源语言而忽视低资源语言。持续预训练(CPT)被视为缓解这一不平衡的可行方法,但单语、双语及代码增强数据策略的有效性尚不明确。本研究系统评估了36种涉及三种多语言基础模型的CPT配置,覆盖30余种语言,按利他、自私、停滞三类划分,涵盖多种资源水平。结果揭示三大发现:(1) 双语CPT虽提升多语言分类表现,但生成时易引发语言混淆;(2) 在CPT中引入编程代码数据可稳定提升多语言分类准确率,尤其利于低资源语言,但导致生成质量轻微下降;(3) 与以往认知相反,语言分类与其对跨语言迁移的影响存在显著偏差:利他型语言常对相关语言产生负面影响,自私型语言行为受配置影响显著,停滞型语言在特定CPT条件下表现出意外适应性。这些复杂互动凸显多语言表征学习的深层挑战,强调需通过系统研究建立可推广的语言分类框架,以指导未来多语言持续预训练策略。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit significant disparities in performance across languages, primarily benefiting high-resource languages while marginalizing underrepresented ones. Continual Pretraining (CPT) has emerged as a promising approach to address this imbalance, although the relative effectiveness of monolingual, bilingual, and code-augmented data strategies remains unclear. This study systematically evaluates 36 CPT configurations involving three multilingual base models, across 30+ languages categorized as altruistic, selfish, and stagnant, spanning various resource levels. Our findings reveal three major insights: (1) Bilingual CPT improves multilingual classification but often causes language mixing issues during generation. (2) Including programming code data during CPT consistently enhances multilingual classification accuracy, particularly benefiting low-resource languages, but introduces a trade-off by slightly degrading generation quality. (3) Contrary to prior work, we observe substantial deviations from language classifications according to their impact on cross-lingual transfer: Languages classified as altruistic often negatively affect related languages, selfish languages show conditional and configuration-dependent behavior, and stagnant languages demonstrate surprising adaptability under certain CPT conditions. These nuanced interactions emphasize the complexity of multilingual representation learning, underscoring the importance of systematic studies on generalizable language classification to inform future multilingual CPT strategies.

多语言持续预训练低资源语言代码数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。