arXiv:2509.09459cs.IR2025-09EMNLP被引 5

通过高质量难负样本提升多语言稠密检索的数据利用效率

Boosting Data Utilization for Multilingual Dense Retrieval

  • 构建难负样本生成机制,提升负样本质量
  • 在16语言的MIRACL数据集上超越多个强基线
  • 适合关注多语言信息检索与数据高效利用的研究者

多语言稠密检索旨在基于统一的检索模型跨语言获取相关文档,其核心挑战在于将不同语言的表示对齐到共享向量空间。现有方法通常通过对比学习微调稠密检索器,但性能高度依赖负样本质量与小批量数据的有效性。不同于以往聚焦复杂模型架构的研究,本文提出一种提升多语言稠密检索中数据利用率的方法,通过生成高质量难负样本并优化小批量数据构成。在包含16种语言的多语言检索基准MIRACL上的大量实验表明,该方法显著优于多个现有强基线。

原文摘要 · Abstract (English)

Multilingual dense retrieval aims to retrieve relevant documents across different languages based on a unified retriever model. The challenge lies in aligning representations of different languages in a shared vector space. The common practice is to fine-tune the dense retriever via contrastive learning, whose effectiveness highly relies on the quality of the negative sample and the efficacy of mini-batch data. Different from the existing studies that focus on developing sophisticated model architecture, we propose a method to boost data utilization for multilingual dense retrieval by obtaining high-quality hard negative samples and effective mini-batch data. The extensive experimental results on a multilingual retrieval benchmark, MIRACL, with 16 languages demonstrate the effectiveness of our method by outperforming several existing strong baselines.

多语言检索稠密检索负样本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。