arXiv:2601.02956cs.CL2026-01ACL被引 1

纠正多语言检索增强生成中的语言偏见,提升低资源语言性能

Enhancing Multilingual RAG Systems with Debiased Language Preference-Guided Query Fusion

  • 提出去偏度量DeLP,消除评估基准的结构偏差
  • 发现检索器本质偏好查询与文档同语种匹配,而非英语优势
  • 设计DELTA框架,通过单语对齐提升跨语言效果,适合多语言应用

多语言检索增强生成(mRAG)系统常表现出对高资源语言(尤其是英语)的感知偏好,导致广泛采用英语中转策略。尽管以往研究归因于大语言模型(LLM)的英语中心能力,但我们发现此类评估结果被评价基准中的结构性先验严重扭曲。具体而言,我们识别出暴露偏差、黄金答案可用性先验(由英语资源集中引发)以及根植于主题局部性的文化先验,均阻碍了对真实语言偏好的准确评估。为此,我们提出去偏语言偏好度量DeLP,显式剔除这些结构混淆因素。使用DeLP分析表明,此前报告的英语偏好主要源于证据分布,并非模型内在偏差。相反,我们发现检索器本质上更倾向查询与文档语言的单语对齐。基于此洞察,我们提出轻量高效的DELTA框架,通过战略性利用单语对齐优化跨语言检索与生成。实验结果表明,DELTA在多种语言上持续优于英语中转和mRAG基线。

原文摘要 · Abstract (English)

Multilingual Retrieval-Augmented Generation (mRAG) systems often exhibit a perceived preference for high-resource languages, particularly English, resulting in the widespread adoption of English pivoting. While prior studies attribute this advantage to the superior English-centric capabilities of Large Language Models (LLMs), we find that such measurements are significantly distorted by structural priors inherent in evaluation benchmarks. Specifically, we identify exposure bias and a gold availability prior-both driven by the disproportionate concentration of resources in English-as well as cultural priors rooted in topic locality, as factors that hinder accurate assessment of genuine language preference. To address these biases, we propose DeLP (Debiased Language Preference), a calibrated metric designed to explicitly factor out these structural confounds. Our analysis using DeLP reveals that the previously reported English preference is largely a byproduct of evidence distribution rather than an inherent model bias. Instead, we find that retrievers fundamentally favor monolingual alignment between the query and the document language. Building on this insight, we introduce DELTA (DEbiased Language preference-guided Text Augmentation), a lightweight and efficient mRAG framework that strategically leverages monolingual alignment to optimize cross-lingual retrieval and generation. Experimental results demonstrate that DELTA consistently outperforms English pivoting and mRAG baselines across diverse languages.

多语言RAG去偏检索生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。