面向欧洲24种官方语言的70亿参数大模型,打破英语主导格局。
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
- 基于60%非英语数据训练,自研多语言分词器提升低资源语言表现。
- 在欧洲版ARC、HellaSwag、TruthfulQA测试中均达领先水平。
- 适合需要支持多语种的欧盟地区应用与研究者使用。
我们推出两款多语言大模型——Teuken 7B-base 和 Teuken 7B-instruct,旨在涵盖欧盟24种官方语言。模型基于约60%非英语数据训练,并采用自研多语言分词器,克服现有大模型以英语为主或仅覆盖少数高资源语言的局限。本文详述模型设计原则,包括数据构成、分词器优化与训练方法。在欧洲版本的ARC、HellaSwag和TruthfulQA基准上,模型表现出色,验证了其跨语言能力。该工作为欧洲本土化大模型提供了可行路径。
原文摘要 · Abstract (English)
We present two multilingual LLMs, Teuken 7B-base and Teuken 7B-instruct, designed to embrace Europe's linguistic diversity by supporting all 24 official languages of the European Union. Trained on a dataset comprising around 60% non-English data and utilizing a custom multilingual tokenizer, our models address the limitations of existing LLMs that predominantly focus on English or a few high-resource languages. We detail the models' development principles, i.e., data composition, tokenizer optimization, and training methodologies. The models demonstrate strong performance across multilingual benchmarks, as evidenced by their performance on European versions of ARC, HellaSwag, and TruthfulQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。