arXiv:2409.16235cs.CL2024-09被引 120

打造欧洲多语言开源大模型,覆盖所有欧盟官方语言。

EuroLLM: Multilingual Language Models for Europe

  • 构建多语言数据集并设计专用分词器,适配欧盟全部官方语言。
  • 发布两个17亿参数模型,支持跨语言理解与生成任务。
  • 适合关注欧洲本地化AI、多语言应用的研究者与开发者。

开源大模型质量显著提升,但仍以英语为主。本文介绍EuroLLM项目,旨在开发一套可开源的多语言大模型,支持所有欧盟官方语言及若干相关语言的文本理解和生成。目前已完成数据收集与过滤、缩放定律建模、多语言分词器设计,以及数据配置与模型架构优化。同时发布初始模型:EuroLLM-1.7B 和 EuroLLM-1.7B-Instruct,评估结果表明其在多语言通用基准和机器翻译任务中表现良好。

原文摘要 · Abstract (English)

The quality of open-weight LLMs has seen significant improvement, yet they remain predominantly focused on English. In this paper, we introduce the EuroLLM project, aimed at developing a suite of open-weight multilingual LLMs capable of understanding and generating text in all official European Union languages, as well as several additional relevant languages. We outline the progress made to date, detailing our data collection and filtering process, the development of scaling laws, the creation of our multilingual tokenizer, and the data mix and modeling configurations. Additionally, we release our initial models: EuroLLM-1.7B and EuroLLM-1.7B-Instruct and report their performance on multilingual general benchmarks and machine translation.

多语言模型开源模型欧盟语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。