发现大模型答事实题时,语言不同答案也不同,可信度不稳定。
Language Models' Factuality Depends on the Language of Inquiry
- 构建1.3万条跨语言事实数据集,量化模型跨语言知识迁移能力。
- 顶尖模型在多语言间知识转移失败率超60%,英语提问准确率显著下降。
- 提示研究者关注语言特定可靠性,推动更鲁棒的多语言模型发展。
多语言大模型本应跨语言一致地回忆事实知识,但常出现知识无法迁移的现象——即使模型在一种语言中掌握正确信息,换用另一种语言提问却频繁出错。例如,模型在阿拉伯语提问下能正确回答拉希德·阿尔沙沙伊来自沙特阿拉伯,但在英语或斯瓦希里语提问时却持续错误。为系统研究此问题,我们构建了一个包含10,000条与国家相关事实、覆盖13种语言的基准测试集,并提出三项新指标:事实召回率(Factual Recall Score)、知识可转移性评分(Knowledge Transferability Score)和跨语言事实知识转移性评分(Cross-Lingual Factual Knowledge Transferability Score),用于评估模型在不同语言中的事实回忆与知识迁移能力。结果揭示当前最先进语言模型在跨语言泛化方面存在根本缺陷,其表现严重依赖提问语言,导致不一致的结果。研究强调模型需具备识别语言特定事实可靠性的能力,并优先利用各语言中最可信的信息。我们已公开该基准与评估框架,以推动未来多语言知识迁移的研究。
原文摘要 · Abstract (English)
Multilingual language models (LMs) are expected to recall factual knowledge consistently across languages, yet they often fail to transfer knowledge between languages even when they possess the correct information in one of the languages. For example, we find that an LM may correctly identify Rashed Al Shashai as being from Saudi Arabia when asked in Arabic, but consistently fails to do so when asked in English or Swahili. To systematically investigate this limitation, we introduce a benchmark of 10,000 country-related facts across 13 languages and propose three novel metrics: Factual Recall Score, Knowledge Transferability Score, and Cross-Lingual Factual Knowledge Transferability Score-to quantify factual recall and knowledge transferability in LMs across different languages. Our results reveal fundamental weaknesses in today's state-of-the-art LMs, particularly in cross-lingual generalization where models fail to transfer knowledge effectively across different languages, leading to inconsistent performance sensitive to the language used. Our findings emphasize the need for LMs to recognize language-specific factual reliability and leverage the most trustworthy information across languages. We release our benchmark and evaluation framework to drive future research in multilingual knowledge transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。