构建首个中文外的印度语检索基准,实现零样本中文检索
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5
- 创建15个数据集的印地语检索基准Hindi-BEIR
- 发现多语言模型在印地语上表现受任务和领域影响
- 提出NLLB-E5模型,无需训练数据即可支持印地语
全球有大量印地语使用者,亟需高效可靠的印地语信息检索系统。尽管已有研究,但针对印地语检索模型的综合性评估基准仍严重缺乏。为此,我们提出了Hindi-BEIR基准,包含15个数据集,覆盖七类不同任务。我们在该基准上评估了先进的多语言检索模型,识别出影响印地语检索性能的任务与领域特定挑战。基于这些发现,我们提出了NLLB-E5模型,采用零样本方法,在无需印地语训练数据的情况下支持印地语检索。我们的贡献包括发布Hindi-BEIR基准和NLLB-E5模型,可为研究人员提供宝贵资源,推动多语言检索模型的发展。
原文摘要 · Abstract (English)
Given the large number of Hindi speakers worldwide, there is a pressing need for robust and efficient information retrieval systems for Hindi. Despite ongoing research, comprehensive benchmarks for evaluating retrieval models in Hindi are lacking. To address this gap, we introduce the Hindi-BEIR benchmark, comprising 15 datasets across seven distinct tasks. We evaluate state-of-the-art multilingual retrieval models on the Hindi-BEIR benchmark, identifying task and domain-specific challenges that impact Hindi retrieval performance. Building on the insights from these results, we introduce NLLB-E5, a multilingual retrieval model that leverages a zero-shot approach to support Hindi without the need for Hindi training data. We believe our contributions, which include the release of the Hindi-BEIR benchmark and the NLLB-E5 model, will prove to be a valuable resource for researchers and promote advancements in multilingual retrieval models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。