跨语言评估大模型幻觉率,发现小模型和多语言模型更易出错。
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
- 用翻译+人工标注构建多语言幻觉检测模型,实现30语言大范围评估。
- 高资源语言因回答更长而幻觉总词数多,但幻觉率与数字足迹无关。
- 小模型和多语言支持的模型幻觉率更高,提示训练策略需改进。
在信息虚假泛滥的时代,大语言模型(LLMs)生成非事实或不忠实内容的幻觉现象成为其全球应用的主要风险。尽管LLMs日益具备多语言能力,现有幻觉研究仍以英语为中心,且集中于机器翻译(MT)和摘要任务,这些任务在真实场景中并不常见。本文旨在量化知识密集型长文本问答(LFQA)中多语言背景下LLM的幻觉程度。我们训练了一个多语言幻觉检测模型,并在30种语言和6个开源LLM家族上开展大规模研究。基于英文幻觉检测数据集,通过机器翻译进行模型训练;同时对5种高资源语言进行人工标注黄金数据。结果显示,银数据(LLM生成)与金数据在幻觉率估计上高度一致,验证了银数据可用于其他语言的估算。最终,我们构建了覆盖30语言的开放域问答数据集,使用LLM生成的问题和维基百科文章作为参考。分析表明:虽然高资源语言因响应更长而幻觉总词数更多,但标准化后的幻觉率与语言数字足迹大小无显著相关性。此外,较小的LLM幻觉率更高,且具备更广语言支持的模型幻觉率也显著更高。
原文摘要 · Abstract (English)
In the age of misinformation, hallucination - the tendency of Large Language Models (LLMs) to generate non-factual or unfaithful responses - represents the main risk for their global utility. Despite LLMs becoming increasingly multilingual, the vast majority of research on detecting and quantifying LLM hallucination are (a) English-centric and (b) focus on machine translation (MT) and summarization, tasks that are less common in realistic settings than open information seeking. In contrast, we aim to quantify the extent of LLM hallucination across languages in knowledge-intensive long-form question answering (LFQA). To this end, we train a multilingual hallucination detection model and conduct a large-scale study across 30 languages and 6 open-source LLM families. We start from an English hallucination detection dataset and rely on MT to translate-train a detection model. We also manually annotate gold data for five high-resource languages; we then demonstrate, for these languages, that the estimates of hallucination rates are similar between silver (LLM-generated) and gold test sets, validating the use of silver data for estimating hallucination rates for other languages. For the final rates estimation, we build open-domain QA dataset for 30 languages with LLM-generated prompts and Wikipedia articles as references. Our analysis shows that LLMs, in absolute terms, hallucinate more tokens in high-resource languages due to longer responses, but that the actual hallucination rates (i.e., normalized for length) seems uncorrelated with the sizes of languages' digital footprints. We also find that smaller LLMs hallucinate more, and significantly, LLMs with broader language support display higher hallucination rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。