220亿参数多语言模型,专为欧洲语种设计
EuroLLM-22B: Technical Report
- 从头训练覆盖24个欧盟官方语言及11种附加语言
- 在多语言评测中表现媲美同规模主流模型
- 开源模型、数据与代码,推动欧洲语言AI研究
本报告介绍EuroLLM-22B,一个从零训练的大型语言模型,旨在满足欧洲公民需求,覆盖所有24个欧盟官方语言及11种额外语言。针对现有开源大模型对欧洲语言支持不足的问题,我们全面阐述了EuroLLM-22B的开发过程,包括分词器设计、架构规格、数据筛选及训练流程。在广泛的多语言基准测试中,该模型在推理、指令遵循和翻译任务上均表现出色,性能达到同类规模模型的水平。为促进后续研究,我们公开发布基础模型与指令微调版本、多语言网络预训练数据、更新版EuroBlocks指令数据集,以及预训练和评估代码库。
原文摘要 · Abstract (English)
This report presents EuroLLM-22B, a large language model trained from scratch to support the needs of European citizens by covering all 24 official European Union languages and 11 additional languages. EuroLLM addresses the issue of European languages being underrepresented and underserved in existing open large language models. We provide a comprehensive overview of EuroLLM-22B's development, including tokenizer design, architectural specifications, data filtering, and training procedures. Across a broad set of multilingual benchmarks, EuroLLM-22B demonstrates strong performance in reasoning, instruction following, and translation, achieving results competitive with models of comparable size. To support future research, we release our base and instruction-tuned models, our multilingual web pretraining data and updated EuroBlocks instruction datasets, as well as our pre-training and evaluation codebases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。