arXiv:2605.15081cs.CLcs.AI2026-05中稿 · ICML

ML-Embed打造高效多语言嵌入模型,突破计算与语言覆盖瓶颈。

ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World

  • 采用三层套娃学习框架,提升参数与推理效率
  • 在430项任务中9项刷新MTEB基准,低资源语言表现优异
  • 开源全部模型数据代码,推动可复现的公平AI发展

高质量文本嵌入的发展正面临三大障碍:高昂的计算成本、对多数世界语言的忽视,以及闭源或开权重模型缺乏透明度。为此,我们提出ML-Embed,一套基于全新3维套娃学习(3D-ML)框架的包容性高效模型。该框架通过套娃表示学习(MRL)和套娃层学习(MLL)实现全生命周期效率优化,并引入套娃嵌入学习(MEL)进一步提升参数效率。为解决语言覆盖问题,我们构建了大规模多语言数据集,训练出参数量从140M到8B不等的系列模型。所有模型、数据与代码均公开。在430项任务上的评估表明,其在17个MTEB基准中9项创纪录,尤其在低资源语言上表现突出,为构建全球公平且高效的AI系统提供了可复现范本。

原文摘要 · Abstract (English)

The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitive computational costs, a narrow linguistic focus that neglects most of the world's languages, and a lack of transparency from closed-source or open-weight models that stifles research. To dismantle these barriers, we introduce ML-Embed, a suite of inclusive and efficient models built upon a new framework: 3-Dimensional Matryoshka Learning (3D-ML). Our framework addresses the computational challenge with comprehensive efficiency across the entire model lifecycle. Beyond the storage benefits of Matryoshka Representation Learning (MRL) and flexible inference-time depth provided by Matryoshka Layer Learning (MLL), we introduce Matryoshka Embedding Learning (MEL) for enhanced parameter efficiency. To address the linguistic challenge, we curate a massively multilingual dataset and train a suite of models ranging from 140M to 8B parameters. In a direct commitment to transparency, we release all models, data, and code. Extensive evaluation on 430 tasks demonstrates that our models set new records on 9 of 17 evaluated MTEB benchmarks, with particularly strong results in low-resource languages, providing a reproducible blueprint for building globally equitable and computationally efficient AI systems.

多语言嵌入模型高效计算开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。