EuroBERT 是面向欧洲及全球语言的多语言编码器,支持超长文本处理。
EuroBERT: Scaling Multilingual Encoders for European Languages
- 基于最新进展重设计多语言编码器,提升跨语言能力
- 在多语言、数学、编程任务上优于现有模型,支持8192词元序列
- 适合需要长文本理解与多语言支持的研究者和开发者
通用多语言向量表示通常由双向编码器模型获得,广泛应用于检索、回归和分类任务。尽管应用广泛,编码器近期被生成式解码器模型的进展所掩盖。然而,推动这些进展的许多创新并不天然依赖于解码器。本文从这些新进展的角度重新审视多语言编码器的发展,提出 EuroBERT,一个涵盖欧洲及广泛使用全球语言的多语言编码器家族。我们的模型在多种任务中表现优于现有方法,包括多语言能力、数学推理和代码生成,并原生支持最长达8,192个词元的序列。我们还分析了 EuroBERT 的设计决策,包括数据集构成和训练流程。所有 EuroBERT 模型(含中间训练检查点)及训练框架均已公开发布。
原文摘要 · Abstract (English)
General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models. Despite their wide applicability, encoders have been recently overshadowed by advances in generative decoder-only models. However, many innovations driving this progress are not inherently tied to decoders. In this paper, we revisit the development of multilingual encoders through the lens of these advances, and introduce EuroBERT, a family of multilingual encoders covering European and widely spoken global languages. Our models outperform existing alternatives across a diverse range of tasks, spanning multilingual capabilities, mathematics, and coding, and natively supporting sequences of up to 8,192 tokens. We also examine the design decisions behind EuroBERT, offering insights into our dataset composition and training pipeline. We publicly release the EuroBERT models, including intermediate training checkpoints, together with our training framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。