arXiv:2410.18270cs.CL2024-10被引 6

跨语言幻觉差异大,低资源语言更易出错。

Multilingual Hallucination Gaps in Large Language Models

  • 用事实评分框架扩展至19种语言,量化幻觉率
  • 低资源语言幻觉率显著高于高资源语言
  • 提醒评估多语言生成时需关注语言差异

大型语言模型(LLMs)日益替代传统搜索引擎,因其能生成类人化文本。但此类模型常产生幻觉——看似可信的误导或虚假信息。本研究探索自由文本生成中的多语言幻觉现象,聚焦所谓‘多语言幻觉差距’,即不同语言和提示下幻觉出现频率的差异。通过引入FactScore度量并拓展至多语言场景,我们对来自LLaMA、Qwen和Aya系列的模型在19种语言中生成传记内容的表现进行了评估,并与维基百科页面对比。结果显示,幻觉率在不同语言间存在显著差异,尤其在高资源与低资源语言之间。这揭示了多语言环境下模型表现不均的问题,也凸显了评估多语言自由文本生成中幻觉挑战的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as alternatives to traditional search engines given their capacity to generate text that resembles human language. However, this shift is concerning, as LLMs often generate hallucinations, misleading or false information that appears highly credible. In this study, we explore the phenomenon of hallucinations across multiple languages in freeform text generation, focusing on what we call multilingual hallucination gaps. These gaps reflect differences in the frequency of hallucinated answers depending on the prompt and language used. To quantify such hallucinations, we used the FactScore metric and extended its framework to a multilingual setting. We conducted experiments using LLMs from the LLaMA, Qwen, and Aya families, generating biographies in 19 languages and comparing the results to Wikipedia pages. Our results reveal variations in hallucination rates, especially between high and low resource languages, raising important questions about LLM multilingual performance and the challenges in evaluating hallucinations in multilingual freeform text generation.

幻觉检测多语言评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。