研究低资源语言对话中的幻觉问题,发现中英文幻觉差异显著。
Investigating Hallucination in Conversations for Low Resource Languages
- 对比GPT-3.5等六模型在印地语、波斯语、汉语对话中的幻觉表现。
- 中文幻觉极少,印地语和波斯语幻觉数量显著更高。
- 揭示低资源语言幻觉风险,为多语言模型优化提供依据。
大型语言模型(LLMs)在生成类人文本方面表现出色,但常产生事实性错误,即所谓‘幻觉’。提升模型可靠性需解决幻觉问题。现有研究主要聚焦英语,本研究将分析扩展至印地语、波斯语和汉语的对话数据。我们对GPT-3.5、GPT-4o、Llama-3.1、Gemma-2.0、DeepSeek-R1和Qwen-3六种模型在三种语言上的事实与语言错误进行了系统评估。结果显示,模型在汉语对话中幻觉极少,但在印地语和波斯语中产生的幻觉数量显著更高。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable proficiency in generating text that closely resemble human writing. However, they often generate factually incorrect statements, a problem typically referred to as 'hallucination'. Addressing hallucination is crucial for enhancing the reliability and effectiveness of LLMs. While much research has focused on hallucinations in English, our study extends this investigation to conversational data in three languages: Hindi, Farsi, and Mandarin. We offer a comprehensive analysis of a dataset to examine both factual and linguistic errors in these languages for GPT-3.5, GPT-4o, Llama-3.1, Gemma-2.0, DeepSeek-R1 and Qwen-3. We found that LLMs produce very few hallucinated responses in Mandarin but generate a significantly higher number of hallucinations in Hindi and Farsi.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。