arXiv:2506.04079cs.CLcs.AI2025-06被引 16

欧洲自研90亿参数多语言大模型,覆盖24种欧盟官方语言

EuroLLM-9B: Technical Report

  • 从零训练,专为欧洲语言设计,含24种官方语言及11种附加语言
  • 通过AI过滤器和合成数据提升多语言覆盖,性能达同类开源模型领先水平
  • 开源全部模型与工具,助力学术研究与本地化应用

本文介绍EuroLLM-9B,一个从零训练的大型语言模型,旨在满足欧洲公民需求,涵盖所有24种欧盟官方语言及11种额外语言。该模型解决现有开源大模型中欧洲语言代表性不足的问题。报告全面概述了EuroLLM-9B的开发过程,包括分词器设计、架构规格、数据筛选及训练流程。介绍了预训练数据的收集与筛选管道,包括基于AI的多语言过滤器EuroFilter,以及用于后训练的新型合成数据集EuroBlocks-Synthetic,可增强欧洲语言覆盖。评估结果显示,EuroLLM-9B在多语言基准和机器翻译任务上表现优异,成为同规模下领先的开源欧洲制造大模型。为支持开放研究与应用,本工作公开了主要组件,包括基础模型、指令微调模型、EuroFilter分类器及合成后训练数据集。

原文摘要 · Abstract (English)

This report presents EuroLLM-9B, a large language model trained from scratch to support the needs of European citizens by covering all 24 official European Union languages and 11 additional languages. EuroLLM addresses the issue of European languages being underrepresented and underserved in existing open large language models. We provide a comprehensive overview of EuroLLM-9B's development, including tokenizer design, architectural specifications, data filtering, and training procedures. We describe the pre-training data collection and filtering pipeline, including the creation of EuroFilter, an AI-based multilingual filter, as well as the design of EuroBlocks-Synthetic, a novel synthetic dataset for post-training that enhances language coverage for European languages. Evaluation results demonstrate EuroLLM-9B's competitive performance on multilingual benchmarks and machine translation tasks, establishing it as the leading open European-made LLM of its size. To support open research and adoption, we release all major components of this work, including the base and instruction-tuned models, the EuroFilter classifier, and the synthetic post-training dataset.

多语言大模型欧洲语言开源模型指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。