arXiv:2501.00874cs.CLcs.IR2025-01被引 1

让大模型嵌入模型零样本支持多语言,提升低资源语言表现

LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models

  • 用轻量连接器融合多语言编码器与大模型嵌入器,实现零样本多语言适配
  • 在14种语言、123个数据集上验证,中低资源语言性能显著提升
  • 无需多语言标注数据,适合跨语言检索等实际场景应用

基于大语言模型的嵌入模型虽在英文文本嵌入任务中达到新基准,但主要聚焦英语,多语言能力仍待探索。为此,我们提出LUSIFER,一种无需多语言监督的零样本方法,将大模型嵌入模型适配多语言任务。其架构结合多语言编码器(语言通用学习者)与针对嵌入任务优化的大模型嵌入器,通过少量可训练参数作为连接器,有效传递多语言理解能力。为全面评估性能,我们构建新基准,涵盖5类主要嵌入任务、123个多样化数据集,覆盖14种语言。实验表明,LUSIFER在多种嵌入任务中显著提升多语言表现,尤其对中低资源语言效果明显,且无需显式多语言训练数据。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) based embedding models have established new state-of-the-art benchmarks for text embedding tasks, particularly in dense vector-based retrieval. However, these models predominantly focus on English, leaving multilingual embedding capabilities largely unexplored. To address this limitation, we present LUSIFER, a novel zero-shot approach that adapts LLM-based embedding models for multilingual tasks without requiring multilingual supervision. LUSIFER's architecture combines a multilingual encoder, serving as a language-universal learner, with an LLM-based embedding model optimized for embedding-specific tasks. These components are seamlessly integrated through a minimal set of trainable parameters that act as a connector, effectively transferring the multilingual encoder's language understanding capabilities to the specialized embedding model. Additionally, to comprehensively evaluate multilingual embedding performance, we introduce a new benchmark encompassing 5 primary embedding tasks, 123 diverse datasets, and coverage across 14 languages. Extensive experimental results demonstrate that LUSIFER significantly enhances the multilingual performance across various embedding tasks, particularly for medium and low-resource languages, without requiring explicit multilingual training data.

多语言嵌入大模型零样本语言泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。