提出多语言检索新评估协议,区分语义相关与查询语言偏好。
MLAIRE: Multilingual Language-Aware Information Retrieval Evaluation Protocal

- 构建跨语言平行文本池,分离语义检索与语言偏好
- 引入LPR和Lang-nDCG等新指标,量化语言偏好程度
- 揭示主流模型在语言偏好上的隐藏缺陷,适合检索系统优化
多语言信息检索在真实搜索场景中日益重要,用户常在多语言语料中提出跨语言查询。现有评估主要关注语言无关的语义相关性,对不同语言的相关段落同等对待。然而,检索结果的语言属性也影响实用性:用户更倾向可读且能验证的查询语言结果,语言不匹配会加剧生成式检索系统中答案溯源的复杂性。为此,我们提出MLAIRE——一种多语言语言感知信息检索评估协议,将跨语言语义检索与查询语言偏好解耦。MLAIRE构建包含多语言平行段落的受控测试集,可在等价翻译条件下衡量语义准确率与查询语言偏好。我们提出语言感知指标,包括语言偏好率(LPR)和Lang-nDCG,以及四向分解方法,分离语义错误与语言偏好失败。评估31个稠密、稀疏及后期交互型检索器发现,标准指标掩盖了显著差异:语义强的模型可能返回非查询语言的结果,而语言偏好强的模型可能召回语义相关性较低的内容。
原文摘要 · Abstract (English)
Multilingual Information Retrieval is increasingly important in real-world search settings, where users issue queries over mixed-language corpora. Existing evaluations mainly reward language-agnostic semantic relevance, treating relevant passages equally regardless of language. Yet retrieval utility also depends on the language of the retrieved passages: users may prefer results they can read and verify in the query language, and query--passage language mismatch can complicate downstream grounding and answer verification in Retrieval-Augmented Generation systems. To evaluate this language-aware dimension, we introduce MLAIRE, a Multilingual Language-Aware Information Retrieval Evaluation protocol that disentangles cross-lingual semantic retrieval from query-language preference. MLAIRE constructs controlled pools with parallel passages across languages, enabling measurement of semantic retrieval accuracy and query-language preference when equivalent translations are available. We propose language-aware metrics, including Language Preference Rate (LPR) and Lang-nDCG, together with a 4-way decomposition separating semantic and query-language preference failures. Evaluating 31 dense, sparse, and late-interaction retrievers, we show that standard metrics obscure distinct behaviors: semantically strong retrievers may return correct content in a non-query language, while retrievers with stronger query-language preference may retrieve less semantically relevant passages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。