构建首个多语言幻觉检测基准,评估大模型在11种语言中的事实错误。
Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models
- 基于11种语言的7.8万条事实陈述构建跨语言验证数据集
- 发现话题分布与网络资源量影响幻觉频率,存在语言特异性偏差
- 适合关注多语言AI可靠性、幻觉检测与负责任生成的研究者
生成式AI中大型语言模型(LLMs)的幻觉问题严重影响多语言应用的可信度。现有幻觉检测基准主要聚焦英语及少数通用语言,缺乏对多样化语言情境下模型表现不一致性的评估能力。为此,我们提出Poly-FEVER,一个大规模多语言事实验证基准,专为评估LLM幻觉检测而设计。该数据集包含来自FEVER、Climate-FEVER和SciFact的77,973条标注事实陈述,覆盖11种语言,是首个系统分析跨语言幻觉模式的大规模数据集,可用于评估ChatGPT、LLaMA系列等模型。分析揭示话题分布与网络资源可用性显著影响幻觉频率,发现语言特异性偏差影响模型准确性。Poly-FEVER支持跨语言幻觉检测比较,推动更可靠、包容性强的AI系统发展。数据集已公开,可访问:https://huggingface.co/datasets/HanzhiZhang/Poly-FEVER,助力负责任AI、事实核查与多语言NLP研究。
原文摘要 · Abstract (English)
Hallucinations in generative AI, particularly in Large Language Models (LLMs), pose a significant challenge to the reliability of multilingual applications. Existing benchmarks for hallucination detection focus primarily on English and a few widely spoken languages, lacking the breadth to assess inconsistencies in model performance across diverse linguistic contexts. To address this gap, we introduce Poly-FEVER, a large-scale multilingual fact verification benchmark specifically designed for evaluating hallucination detection in LLMs. Poly-FEVER comprises 77,973 labeled factual claims spanning 11 languages, sourced from FEVER, Climate-FEVER, and SciFact. It provides the first large-scale dataset tailored for analyzing hallucination patterns across languages, enabling systematic evaluation of LLMs such as ChatGPT and the LLaMA series. Our analysis reveals how topic distribution and web resource availability influence hallucination frequency, uncovering language-specific biases that impact model accuracy. By offering a multilingual benchmark for fact verification, Poly-FEVER facilitates cross-linguistic comparisons of hallucination detection and contributes to the development of more reliable, language-inclusive AI systems. The dataset is publicly available to advance research in responsible AI, fact-checking methodologies, and multilingual NLP, promoting greater transparency and robustness in LLM performance. The proposed Poly-FEVER is available at: https://huggingface.co/datasets/HanzhiZhang/Poly-FEVER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。