arXiv:2604.07123cs.CL2026-04

多语言大模型在冲突信息中偏好特定语言,可能误导用户判断。

Language Bias under Conflicting Information in Multilingual LLMs

  • 用多语言新闻数据测试模型对冲突信息的响应,发现模型常忽略矛盾
  • 普遍偏向中文,排斥俄语,且更信任与提示语同语言的信息
  • 无论训练地如何,模型都存在语言偏好,适合关注公平性的开发者参考

大型语言模型在处理冲突信息时存在偏差。本文将‘针堆中的针’范式拓展至多语言场景,使用五种语言的真实新闻数据,对不同规模的多语言LLM进行评估。结果表明,所有测试模型(包括GPT-5.2)在绝大多数情况下忽略冲突,自信地仅选择一个答案。模型普遍存在语言偏好:普遍排斥俄语,长上下文下倾向中文;该偏好在中外训练模型间一致,但中国境内训练模型表现更强。此外,模型倾向于优先采纳与提示语语言一致的信息。研究提醒开发者和用户注意此类语言偏见,推动其成因分析与缓解方法研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been shown to contain biases in the process of integrating conflicting information when answering questions. Here we ask whether such biases also exist with respect to which language is used for each conflicting piece of information. To answer this question, we extend the conflicting needles in a haystack paradigm to a multilingual setting and perform a comprehensive set of evaluations with naturalistic news domain data in five different languages, for a range of multilingual LLMs of different sizes. We find that all LLMs tested, including GPT-5.2, ignore the conflict and confidently assert only one of the possible answers in the large majority of cases. Furthermore, there is a consistent bias across models and prompting languages in which languages are preferred, with a general bias against Russian and, for the longest context lengths, in favor of Chinese. The language preferences are consistent between models trained inside and outside of mainland China, though somewhat stronger in the former category. There is also a general tendency among models to prioritize information that matches the language used for prompting. We hope to make users and developers of multilingual LLMs aware of this category of biases, to spur further research on their causes and possible mitigation.

多语言模型语言偏见信息冲突大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。