多语言大模型在跨语言知识冲突中,资源多的语言更占优,但亲缘语言能逆袭。
When Abundance Conceals Weakness: Knowledge Conflict in Multilingual Models
- 构建四阶段冲突解决框架,系统评估多语言模型如何应对矛盾信息。
- 在推理任务中,高资源语言更具说服力;在事实类冲突中,语言亲缘性更重要。
- 覆盖10种语言的全新评测集,揭示模型决策受语言资源与亲缘关系双重影响。
大型语言模型在多语言环境下蕴含广泛世界知识,但其内部信念在不同语言间分布不均。当外部证据与模型内存储的记忆冲突时,会产生跨语言知识冲突,这一现象此前主要局限于以英语为中心的研究。本文提出CLEAR(Cross-Lingual Knowledge conflict Evaluation Framework),系统分析多语言大模型如何调和内部信念与多语言外部证据之间的矛盾。CLEAR将冲突解决分解为四个渐进阶段,从多语言参数化诱导到竞争性多源跨语言归纳,并在两个互补的问答基准上评估模型表现。我们构建了涵盖10种类型多样语言的ConflictQA和ConflictingQA多语言版本,评估了六种代表性多语言LLM。实验发现:在需要推理的任务中,冲突解决由语言资源丰度主导,高资源语言更具说服力;而在以实体为中心的事实冲突中,语言亲缘性起决定作用,低资源但语言相近的语言反而优于远距离的高资源语言。
原文摘要 · Abstract (English)
Large Language Models (LLMs) encode vast world knowledge across multiple languages, yet their internal beliefs are often unevenly distributed across linguistic spaces. When external evidence contradicts these language-dependent memories, models encounter \emph{cross-lingual knowledge conflict}, a phenomenon largely unexplored beyond English-centric settings. We introduce \textbf{CLEAR}, a \textbf{C}ross-\textbf{L}ingual knowl\textbf{E}dge conflict ev\textbf{A}luation f\textbf{R}amework that systematically examines how multilingual LLMs reconcile conflicting internal beliefs and multilingual external evidence. CLEAR decomposes conflict resolution into four progressive scenarios, from multilingual parametric elicitation to competitive multi-source cross-lingual induction, and systematically evaluates model behavior across two complementary QA benchmarks with distinct task characteristics. We construct multilingual versions of ConflictQA and ConflictingQA covering 10 typologically diverse languages and evaluate six representative LLMs. Our experiments reveal a task-dependent decision dichotomy. In reasoning-intensive tasks, conflict resolution is dominated by language resource abundance, with high-resource languages exerting stronger persuasive power. In contrast, for entity-centric factual conflicts, linguistic affinity, not resource scale, becomes decisive, allowing low-resource but linguistically aligned languages to outperform distant high-resource ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。