arXiv:2509.21104cs.CL2025-09

首个针对波斯语的幻觉评估基准,揭示大模型在波斯语内容上的幻觉问题。

PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models

  • 构建三阶段LLM生成+人工验证的动态评估流程,聚焦波斯语外在与内在幻觉。
  • 12个模型测试显示普遍难以识别波斯语幻觉,外部知识可部分缓解问题。
  • 专为波斯文化设计,适合研究低资源语言幻觉与本地化模型评估者。

幻觉是影响所有大型语言模型(LLMs)的持续性问题,尤其在低资源语言如波斯语中更为突出。PerHalluEval(波斯语幻觉评估)是首个专为波斯语设计的动态幻觉评估基准。该基准采用三阶段LLM驱动流程,结合人工验证,生成关于问答与摘要任务的合理答案和摘要,重点检测外在与内在幻觉。同时,利用生成标记的对数概率筛选最具可信度的幻觉实例。此外,我们邀请人工标注员在问答数据集中标注波斯语特有语境,以评估模型在波斯文化相关内容上的表现。对12个大语言模型(包括开源与闭源模型)的评估表明,这些模型普遍难以识别波斯语中的幻觉文本。研究发现,提供外部知识(如摘要任务的原始文档)可部分缓解幻觉现象。此外,在幻觉表现上,专为波斯语训练的模型与其他模型无显著差异。

原文摘要 · Abstract (English)

Hallucination is a persistent issue affecting all large language Models (LLMs), particularly within low-resource languages such as Persian. PerHalluEval (Persian Hallucination Evaluation) is the first dynamic hallucination evaluation benchmark tailored for the Persian language. Our benchmark leverages a three-stage LLM-driven pipeline, augmented with human validation, to generate plausible answers and summaries regarding QA and summarization tasks, focusing on detecting extrinsic and intrinsic hallucinations. Moreover, we used the log probabilities of generated tokens to select the most believable hallucinated instances. In addition, we engaged human annotators to highlight Persian-specific contexts in the QA dataset in order to evaluate LLMs' performance on content specifically related to Persian culture. Our evaluation of 12 LLMs, including open- and closed-source models using PerHalluEval, revealed that the models generally struggle in detecting hallucinated Persian text. We showed that providing external knowledge, i.e., the original document for the summarization task, could mitigate hallucination partially. Furthermore, there was no significant difference in terms of hallucination when comparing LLMs specifically trained for Persian with others.

幻觉评估波斯语低资源语言LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。