用单语知识库实现跨语言检索,提升低资源语言信息获取能力
Multilingual Information Retrieval with a Monolingual Knowledge Base
- 通过加权采样优化对比学习,微调多语言嵌入模型
- 在MRR上提升31.03%,Recall@3提升33.98%
- 不依赖多语知识库,适用于多语言及混合语言场景
多语言信息检索已成为促进跨语言知识共享的重要工具。然而,高质量知识库资源往往稀缺且局限于少数语言,因此将不同语言的句子映射到与知识库同语言的特征空间,成为跨语言知识共享的关键。本文提出一种新的微调策略,通过加权采样进行对比学习,实现基于单语知识库的多语言信息检索。实验表明,该方法相比标准策略,在MRR上最高提升31.03%,Recall@3最高提升33.98%。所提方法具有语言无关性,适用于多语言及代码切换场景。
原文摘要 · Abstract (English)
Multilingual information retrieval has emerged as powerful tools for expanding knowledge sharing across languages. On the other hand, resources on high quality knowledge base are often scarce and in limited languages, therefore an effective embedding model to transform sentences from different languages into a feature vector space same as the knowledge base language becomes the key ingredient for cross language knowledge sharing, especially to transfer knowledge available in high-resource languages to low-resource ones. In this paper we propose a novel strategy to fine-tune multilingual embedding models with weighted sampling for contrastive learning, enabling multilingual information retrieval with a monolingual knowledge base. We demonstrate that the weighted sampling strategy produces performance gains compared to standard ones by up to 31.03\% in MRR and up to 33.98\% in Recall@3. Additionally, our proposed methodology is language agnostic and applicable for both multilingual and code switching use cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。