模型回答多语言问题时,会优先选英文资料,哪怕它不最相关。
Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
- 通过分析模型内部机制,控制其他变量研究语言偏好
- 英语查询下模型更倾向引用英文文档,低资源语言偏差更大
- 有时为选母语文档牺牲内容相关性,适合关注多语言公平性的研究者
多语言检索增强生成(mRAG)系统使语言模型能跨语言回答知识密集型问题并提供引用支持。尽管应用日益广泛,一个关键问题是不同语言文档的混合是否会导致生成与引用行为出现非预期偏差。为此,我们提出一种受控方法,利用模型内部信息在保持文档相关性等条件不变的情况下测量语言偏好。在八种语言和六种开源模型上,我们发现当查询为英语时,模型更倾向于引用英语来源,且该偏差在低资源语言及上下文中间位置的文档中更为显著。更重要的是,模型有时会牺牲文档相关性以满足语言偏好,表明引用选择并非始终由信息量驱动。研究揭示了语言模型如何利用多语言上下文并影响引用行为。
原文摘要 · Abstract (English)
Multilingual Retrieval-Augmented Generation (mRAG) systems enable language models to answer knowledge-intensive queries with citation-supported responses across languages. Despite their growing use, an open questions is whether the mixture of different document languages impacts generation and citation behavior in unintended ways. To investigate this, we introduce a controlled methodology using model internals to measure language preference while holding other factors such as document relevance constant. Across eight languages and six open-weight models, we find that models preferentially cite English sources when queries are in English, with this bias amplified for lower-resource languages and for documents positioned mid-context. More crucially, we find that models sometimes trade-off document relevance for language preference, indicating that citation choices are not always driven by informativeness alone. Our findings shed light on how language models leverage multilingual context and influence citation behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。