arXiv:2605.13521cs.IR2026-05被引 3

多语言嵌入模型支持200+语言,长文本检索效果领先。

Granite Embedding Multilingual R2 Models

论文配图:Granite Embedding Multilingual R2 Models
图 1 · 摘自论文原文
  • 基于ModernBERT架构,支持52种语言和代码的双编码器模型。
  • 32,768令牌上下文窗口,性能在多语言检索中达到顶尖水平。
  • 开源可商用,适合企业级应用与负责任的研究使用。

我们推出了多语言Granite Embedding R2模型系列,是一组面向200多种语言的企业级稠密检索编码器模型。在之前的英文R2基础上,新增对52种语言和编程代码的支持,上下文窗口扩展至32,768个令牌(相比R1提升64倍),在多语言、跨语言文本搜索、代码检索、长文档搜索及推理检索数据集上均达到当前最优表现。该系列包含两个基于ModernBERT架构的双编码器模型:一个311M参数的全尺寸模型,以及一个通过模型剪枝和词表筛选构建的97M参数紧凑模型,后者在小于100M参数的开源多语言嵌入模型中取得了最高检索分数。全尺寸模型还支持马特约什卡表示学习,实现嵌入维度灵活调整。所有模型均在符合企业规范的数据上训练,并有治理监督,以Apache 2.0许可证发布于HuggingFace,旨在支持负责任使用并促进科研与企业自由采用。

原文摘要 · Abstract (English)

We introduce the multilingual Granite Embedding R2 models, a family of encoder-based embedding models for enterprise-scale dense retrieval across 200+ languages. Extending our English-focused R2 release, these models add enhanced support for 52 languages and programming code, a 32,768-token context window (a 64x expansion over R1), and state-of-the-art overall performance across multilingual and cross-lingual text search, code retrieval, long-document search, and reasoning retrieval datasets. The release consists of two bi-encoder models based on the ModernBERT architecture with an expanded multilingual vocabulary: a 311M-parameter full-size, and a 97M-parameter compact model built via model pruning and vocabulary selection that achieves the highest retrieval score of any open multilingual embedding model under 100M parameters. The full-size also supports Matryoshka Representation Learning for flexible embedding dimensionality. Both models are trained on enterprise-appropriate data with governance oversight, and released under the Apache 2.0 license at https://huggingface.co/collections/ibm-granite, designed to support responsible use and enable unrestricted research and enterprise adoption.

多语言嵌入稠密检索代码检索开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。