arXiv:2509.05668cs.CLcs.AI2025-09被引 5

Llama-GENBA-10B是首个融合德语、英语和巴伐利亚语的100亿参数大模型,缓解英语主导问题。

Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian

  • 基于Llama 3.1-8B扩展至100亿参数,三语持续预训练共1640亿词元
  • 在巴伐利亚语任务上超越Apertus-8B-2509和gemma-2-9b,德语性能媲美EuroLLM
  • 首次建立三语评估基准,助力低资源语言模型发展

我们提出Llama-GENBA-10B,一个针对英语主导偏见的三语基础模型。该模型在Llama 3.1-8B基础上扩展至100亿参数,持续预训练于1640亿词元(820亿英语、820亿德语、8000万巴伐利亚语),在资源分配上实现平衡并防止英语主导。面向德语自然语言处理社区,同时推动巴伐利亚语这一低资源语言的发展。开发过程中解决四大挑战:(1) 尽管巴伐利亚语稀缺,仍构建多语种语料库;(2) 设计统一分词器支持英语、德语与巴伐利亚语;(3) 优化架构及语言比例超参数以提升跨语言迁移能力;(4) 首次将德语基准翻译为巴伐利亚语,建立标准化三语评估套件。评估显示,微调后的模型在巴伐利亚语任务上优于Apertus-8B-2509和gemma-2-9b,成为同类最佳;在英语上超越EuroLLM,德语表现与之持平。在Cerebras CS-2上训练展示了高效的大规模多语言预训练,并记录了能耗数据,为包容性基础模型提供了可复用范式。

原文摘要 · Abstract (English)

We present Llama-GENBA-10B, a trilingual foundation model addressing English-centric bias in large language models. Built on Llama 3.1-8B and scaled to 10B parameters, Llama-GENBA-10B is continuously pretrained on 164B tokens (82B English, 82B German, and 80M Bavarian), balancing resources while preventing English dominance. Targeted at the German NLP community, the model also promotes Bavarian as a low-resource language. Development tackled four challenges: (1) curating a multilingual corpus despite Bavarian scarcity, (2) creating a unified tokenizer for English, German, and Bavarian, (3) optimizing architecture and language-ratio hyperparameters for cross-lingual transfer, and (4) establishing the first standardized trilingual evaluation suite by translating German benchmarks into Bavarian. Evaluations show that Llama-GENBA-10B achieves strong cross-lingual performance, with the fine-tuned variant surpassing Apertus-8B-2509 and gemma-2-9b in Bavarian and establishing itself as the best model in its class for this language, while also outperforming EuroLLM in English and matching its results in German. Training on the Cerebras CS-2 demonstrated efficient large-scale multilingual pretraining with documented energy use, offering a blueprint for inclusive foundation models that integrate low-resource languages.

三语模型低资源语言巴伐利亚语多语言预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。