arXiv:2410.07809cs.CLcs.LG2024-10被引 1

选语言没标准答案,多语言训练反而可能变差。

No Optimal Language Set Exists for Multilingual Instruction Tuning: Insights from a Linguistically-Informed Study

  • 按语言类型、地理分布等特征选语言,对比随机选择
  • 模型和任务不同,最佳语言组合也不同,加语言会恶化性能
  • 适合关注多语言训练数据设计的研究者和实践者

多语言指令微调(MIT)面临多语言诅咒、数据稀缺和高计算成本的挑战。一个自然假设是:精心挑选语言多样性高的语言集能提升通用性能。我们系统评估了基于语言类型、地理分布、语义和学习特征的选语言策略,对比随机基线,在三种模型(mGPT、mT5-xl、BLOOM)和五个多语言基准上进行测试。关键发现是:在固定预算下,不存在普适最优的语言选择策略——性能高度依赖任务与模型;超过适度语言数量后,性能反而下降,符合多语言诅咒现象。研究讨论了对数据构建的影响,以及基准平均评价的潜在陷阱。所有资源已公开于 https://github.com/GGLAB-KU/ling-informed-mit。

原文摘要 · Abstract (English)

Multilingual instruction tuning (MIT) is challenged by the curse of multilinguality, data scarcity, and high computational cost. A natural hypothesis is that carefully selecting a linguistically diverse set of languages yields universally better models. We test this systematically by evaluating linguistically-informed selection strategies---based on typological, geographical, semantic, and learned features---against random baselines across three model families (mGPT, mT5-xl, BLOOM) and five multilingual benchmarks. Our key negative finding is that no universal language selection strategy emerges in our fixed-budget setting: performance is strongly task- and model-dependent, and adding more languages beyond a modest threshold triggers degradation consistent with the curse of multilinguality. We discuss implications for MIT data curation and the pitfalls of benchmark-averaged evaluation. All resources are publicly available at https://github.com/GGLAB-KU/ling-informed-mit.

多语言指令微调数据选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。