arXiv:2601.12555cs.CL2026-01

测试多语言大模型在上下文中的事实回忆能力,发现上下文会显著降低准确率。

Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models

  • 设计受控提示,让实体通过上下文间接提及而非直接命名
  • 多语言测试显示上下文使事实召回率普遍下降,关系间差异大
  • 模型越大越抗干扰,真实姓名影响不一致,适合评估模型理解力

大语言模型可在多种语言中回忆广泛的事实知识。然而,现有评估主要针对孤立场景下的事实检索——即实体被明确命名且问题直接询问。在自然语言使用中,事实常通过上下文间接获取,目标实体仅隐含出现。本文研究上下文中介的事实回忆,考察当目标实体嵌入自然语境而非被直接查询时,多语言大模型能否可靠地召回事实知识。我们构建了保持原始事实不变但引入指代中介的受控提示,并通过合成名与真实名在五种语言中的对比,分离上下文影响与名称特异性关联。评估多个模型家族发现,上下文中介会持续降低事实回忆表现,不同关系间差异显著;更大模型对上下文更具鲁棒性,性能下降幅度更小;而真实姓名及其来源的影响则混合且无系统规律。结果揭示了孤立事实回忆与依赖上下文的语言理解之间存在差距。

原文摘要 · Abstract (English)

Large language models (LLMs) can recall a wide range of factual knowledge across languages. However, existing factual recall evaluations primarily assess fact retrieval in isolation, where the queried entity is explicitly named and the fact is requested directly. In natural language use, facts are often accessed through context, where the relevant entity is introduced only indirectly. In this work, we study contextually mediated factual recall, asking whether LLMs can reliably retrieve factual knowledge when the target entity is embedded in a naturalistic context rather than queried explicitly, across languages. We construct controlled prompts that preserve the underlying fact while introducing referential mediation through contextual sentences. To disentangle contextual effects from name-specific associations, we further compare performance using synthetic names and real names across languages. Evaluating multiple model families in five languages, we find that contextual mediation consistently degrades factual recall, with substantial variation across relations. Larger models are more robust to contextual mediation, exhibiting a reduced performance gap relative to direct queries, while the effect of real names and name origin is mixed and unsystematic. These findings highlight a gap between isolated factual recall and context-dependent language understanding in multilingual LLMs.

多语言事实召回上下文理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。