arXiv:2501.05629cs.CLcs.AI2025-01中稿 · AAAI被引 1

研究大模型在多语言任务中的表现差异,发现规模对零样本性能影响小,但少样本下大模型显著提升。

The Impact of Model Scaling on Seen and Unseen Language Performance

  • 对比204种语言的零样本与少样本表现,分析三类模型的缩放规律
  • 零样本下性能几乎不随模型变大而提升,两样本中大模型线性增益
  • 资源总量比预训练语言比例更能预测模型多语言能力,适合跨语言研究者

大型语言模型(LLM)在多语言语料上快速演进,迫切需要深入理解其在不同语言和模型规模下的表现。本研究系统考察了三类不同规模模型在204种语言上的文本分类与机器翻译任务中,对已见与未见语言的表现及缩放行为。结果表明,零样本与两样本场景间存在显著差异:零样本下模型规模对性能影响极小,性能基本持平;而在两样本设置中,大模型在多语言文本分类任务中表现出明确的线性提升。翻译任务中,仅指令微调模型受益于规模扩展。分析还显示,整体资源水平(而非预训练语言比例)是更优的性能预测因子,揭示了多语言模型有效性的关键驱动因素。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Models (LLMs), particularly those trained on multilingual corpora, has intensified the need for a deeper understanding of their performance across a diverse range of languages and model sizes. Our research addresses this critical need by studying the performance and scaling behavior of multilingual LLMs in text classification and machine translation tasks across 204 languages. We systematically examine both seen and unseen languages across three model families of varying sizes in zero-shot and few-shot settings. Our findings show significant differences in scaling behavior between zero-shot and two-shot scenarios, with striking disparities in performance between seen and unseen languages. Model scale has little effect on zero-shot performance, which remains mostly flat. However, in two-shot settings, larger models show clear linear improvements in multilingual text classification. For translation tasks, however, only the instruction-tuned model showed clear benefits from scaling. Our analysis also suggests that overall resource levels, not just the proportions of pretraining languages, are better predictors of model performance, shedding light on what drives multilingual LLM effectiveness.

多语言模型缩放少样本学习语言泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。