揭示多语言检索中的语言偏见并提出新方法缓解
Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods
- 以语义等价查询检验不同语言排名一致性
- 发现传统与神经检索方法均存在显著语言偏差
- 提出新损失函数LaKDA,提升多语言公平性
多语言信息检索(MLIR)中的语言公平性对保障多元语言用户平等获取信息至关重要。本文基于不同语言但语义相同的查询在相同多语言文档上应产生相似排序列表的假设,评估了语言公平性。实验涵盖传统检索方法及基于mBERT和XLM-R的DPR神经排序器。结果揭示当前MLIR技术存在内在语言偏见,且不同方法间差异明显。为此,本文提出新型损失函数LaKDA,有效缓解神经MLIR中的语言偏差,显著提升跨语言公平性。
原文摘要 · Abstract (English)
Language fairness in multilingual information retrieval (MLIR) systems is crucial for ensuring equitable access to information across diverse languages. This paper sheds light on the issue, based on the assumption that queries in different languages, but with identical semantics, should yield equivalent ranking lists when retrieving on the same multilingual documents. We evaluate the degree of fairness using both traditional retrieval methods, and a DPR neural ranker based on mBERT and XLM-R. Additionally, we introduce `LaKDA', a novel loss designed to mitigate language biases in neural MLIR approaches. Our analysis exposes intrinsic language biases in current MLIR technologies, with notable disparities across the retrieval methods, and the effectiveness of LaKDA in enhancing language fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。