arXiv:2604.05684cs.IR2026-04被引 4

解决多语言检索中模型偏爱英文文档的问题

Improving Semantic Proximity in Information Retrieval through Cross-Lingual Alignment

  • 设计新评估场景与指标,精准衡量跨语言对齐能力
  • 用2.8k样本训练,显著提升跨语言检索效果
  • 适合关注多语言系统公平性与性能的研究者

随着多语言文档的广泛使用,跨语言信息检索(CLIR)成为重要研究方向。传统设定下,文档语言与查询语言不同,且文档通常为单一语言。本文指出,在英语文档与其他语言共存的场景中,现有多语言检索模型常优先返回无关的英语文档,而非与查询同语种的相关文档。为此,我们设计多种评估场景与指标,系统分析该现象。进一步提出一种新颖的训练策略,仅用2.8k样本即可显著提升跨语言对齐能力,有效缓解英语偏好问题。大量实验表明,该方法能大幅提升多数多语言嵌入模型的跨语言对齐性能。

原文摘要 · Abstract (English)

With the increasing accessibility and utilization of multilingual documents, Cross-Lingual Information Retrieval (CLIR) has emerged as an important research area. Conventionally, CLIR tasks have been conducted under settings where the language of documents differs from that of queries, and typically, the documents are composed in a single coherent language. In this paper, we highlight that in such a setting, the cross-lingual alignment capability may not be evaluated adequately. Specifically, we observe that, in a document pool where English documents coexist with another language, most multilingual retrievers tend to prioritize unrelated English documents over the related document written in the same language as the query. To rigorously analyze and quantify this phenomenon, we introduce various scenarios and metrics designed to evaluate the cross-lingual alignment performance of multilingual retrieval models. Furthermore, to improve cross-lingual performance under these challenging conditions, we propose a novel training strategy aimed at enhancing cross-lingual alignment. Using only a small dataset consisting of 2.8k samples, our method significantly improves the cross-lingual retrieval performance while simultaneously mitigating the English inclination problem. Extensive analyses demonstrate that the proposed method substantially enhances the cross-lingual alignment capabilities of most multilingual embedding models.

跨语言检索多语言模型嵌入对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。