arXiv:2412.04506cs.CLcs.IR2024-12被引 76

多语言检索模型北极嵌入2.0,兼顾英语与多语种性能。

Arctic-Embed 2.0: Multilingual Retrieval Without Compromise

  • 采用新训练方法实现多语言与英语检索性能双优。
  • 支持马特罗什卡表示学习,压缩后质量下降更小。
  • 开源模型,适合需要高效多语检索的开发者。

本文介绍了 Arctic-Embed 2.0 的训练方法,这是一组面向精准高效多语言检索的开源文本嵌入模型。先前工作常因英语检索性能下降而受限,而 Arctic-Embed 2.0 在多语言及仅英语基准上均达到竞争力表现,并支持马特罗什卡表示学习(MRL),相比其他方案显著降低压缩后的质量损失。文中详述了设计与实现,探讨了开发过程中出现的若干关键开放研究问题,通过实验分析并展开深入讨论,旨在推动该领域进一步发展。

原文摘要 · Abstract (English)

This paper presents the training methodology of Arctic-Embed 2.0, a set of open-source text embedding models built for accurate and efficient multilingual retrieval. While prior works have suffered from degraded English retrieval quality, Arctic-Embed 2.0 delivers competitive retrieval quality on multilingual and English-only benchmarks, and supports Matryoshka Representation Learning (MRL) for efficient embedding storage with significantly lower compressed quality degradation compared to alternatives. We detail the design and implementation, presenting several important open research questions that arose during model development. We conduct experiments exploring these research questions and include extensive discussion aimed at fostering further discussion in this field.

多语言嵌入模型检索开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。