arXiv:2604.20199cs.CL2026-04ACL被引 2

解决多语言问答中对非英语的系统性偏见问题

All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG

论文配图:All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG
图 1 · 摘自论文原文
  • 设计无语言依赖的重排序机制,让多语言证据更公平参与
  • 实验显示非英语表现提升显著,最高达18.7%的准确率改善
  • 适合需要公平多语言支持的AI系统开发者使用

多语言检索增强生成(mRAG)利用跨语言知识来增强大模型的全球知识能力。然而我们发现,当前mRAG系统在重排序阶段存在语言偏见,系统性地偏好英语及查询母语。通过引入估计最优基准分析,我们量化了现有重排序器与可实现上限之间的显著性能差距。进一步分析揭示关键分布失配:最优预测需依赖多语言分散的证据,但当前系统却系统性抑制这些‘答案关键’文档,从而限制生成效果。为此,我们提出语言无关的、以生成效用驱动的重排序对齐方法LAURA,使多语言证据排名与下游生成目标一致。在多种语言和生成模型上的实验表明,LAURA有效缓解语言偏见,并持续提升mRAG性能。

原文摘要 · Abstract (English)

Multilingual Retrieval-Augmented Generation (mRAG) leverages cross-lingual evidence to ground Large Language Models (LLMs) in global knowledge. However, we show that current mRAG systems suffer from a language bias during reranking, systematically favoring English and the query's native language. By introducing an estimated oracle evidence analysis, we quantify a substantial performance gap between existing rerankers and the achievable upper bound. Further analysis reveals a critical distributional mismatch: while optimal predictions require evidence scattered across multiple languages, current systems systematically suppress such ``answer-critical'' documents, thereby limiting downstream generation performance. To bridge this gap, we propose \textit{\textbf{L}anguage-\textbf{A}gnostic \textbf{U}tility-driven \textbf{R}eranker \textbf{A}lignment (LAURA)}, which aligns multilingual evidence ranking with downstream generative utility. Experiments across diverse languages and generation models show that LAURA effectively mitigates language bias and consistently improves mRAG performance.

多语言RAG偏见缓解生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。