arXiv:2603.08182cs.CLcs.AI2026-03被引 2

300亿参数模型提升欧洲小语种表现,让语言更平等

TildeOpen LLM: Leveraging Curriculum Learning to Achieve Equitable Language Representation

  • 用课程学习策略平衡数据分布,交替使用均匀与真实语料
  • 在仅用较少算力下超越现有开源多语言模型,尤其改善波罗的海等语种
  • 全开源可下载,适合关注语言公平与低资源语种研究者

大型语言模型因英语等高资源语言主导训练数据,在许多欧洲语言上表现不佳。本文提出 TildeOpen LLM,一个针对34种欧洲语言训练的300亿参数开源基础模型,旨在促进语言公平并提升低资源语言性能。为应对数据不平衡问题,采用数据集上采样结合课程学习训练策略,交替使用均匀与自然语言分布。尽管训练所用计算资源显著更少,该模型在多项多语言基准测试中表现优于其他开源多语言模型,尤其在波罗的海、芬兰-乌戈尔及斯拉夫语系中表现突出。人工评估显示,其语言错误比领先基线减少多达十倍。模型及相关资源已完全开源,可在 huggingface.co/TildeAI/TildeOpen-30b 公开获取。结果表明,通过精心的数据整理与均衡训练策略,可在不增加模型规模或训练量的前提下显著提升多语言模型质量。

原文摘要 · Abstract (English)

Large language models often underperform in many European languages due to the dominance of English and a few high-resource languages in training data. This paper presents TildeOpen LLM, a 30-billion-parameter open-weight foundational model trained for 34 European languages to promote linguistic equity and improve performance for low-resource languages. To address the data imbalance, we combine dataset upsampling with a curriculum-based training schedule that alternates between uniform and natural language distributions. The resulting model performs favorably compared to other multilingual LLMs despite being trained with significantly fewer computing resources. Evaluation across multiple multilingual benchmarks shows that TildeOpen surpasses existing open-weight models in text generation and comprehension, particularly for Baltic, Finno-Ugric, and Slavic languages. Human evaluations confirm an up to tenfold reduction in linguistic errors relative to leading baselines. The model and associated resources are fully open-weight and publicly available at huggingface.co/TildeAI/TildeOpen-30b. These outcomes demonstrate that careful data curation and balanced training strategies can substantially enhance multilingual model quality without increasing model size or training volume.

多语言模型语言公平开源模型课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。