arXiv:2605.17007cs.CL2026-05被引 2

首个针对阿拉伯语大模型幻觉的评测基准,覆盖多类推理与文化场景。

HalluScore: Large Language Model Hallucination Question Answering Benchmark

论文配图:HalluScore: Large Language Model Hallucination Question Answering Benchmark
图 1 · 摘自论文原文
  • 构建结构化阿拉伯语问答数据集,聚焦幻觉检测与评估。
  • 包含827个经验证的高质量问题,覆盖历史、文化、逻辑等多维度。
  • 适用于研究阿拉伯语模型可靠性与文化适配性的学者与开发者。

大型语言模型(LLMs)在自然语言生成方面取得显著进展,但仍易产生幻觉。尽管已有多个英文和中文幻觉评测基准,但阿拉伯语因标注资源稀缺及语言形态复杂,仍缺乏有效评测工具。为此,我们提出 HalluScore——一个结构化的阿拉伯语问答基准,用于评估不同推理难度、知识领域、历史时间线及文化背景下的幻觉行为。该数据集包含827个精心筛选的问题,每个问题均配有真实证据、答案解释和多标签标注。通过该基准,我们对17个阿拉伯语、多语言及推理型大模型进行了全面实证分析,并提供高质量人工标注,识别各模型输出中的幻觉、非幻觉及部分幻觉响应。结果表明,阿拉伯语模型的幻觉不仅限于事实错误,还涉及文化理解、语言推理与逻辑一致性等深层挑战。我们已公开 HalluScore,以支持未来提升阿拉伯语大模型可靠性和文化适应性的研究。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable progress in natural language generation, but remain susceptible to hallucination. In response to growing concerns about hallucinations, several benchmarks have been developed, primarily in English and Chinese. However, Arabic remains underrepresented, with limited benchmarks for LLMs hallucination due to scarce annotated resources and the language's morphological complexity. Consequently, existing benchmarks do not adequately reflect the linguistic, cultural, and reasoning characteristics of Arabic. To address this gap, we introduce HalluScore, a structured Arabic question answering benchmark designed to evaluate hallucination behavior in LLMs across different levels of reasoning difficulty, various knowledge domains, historical timelines, and culturally grounded Arabic scenarios. It contains 827 carefully curated questions for evaluating, detecting, and mitigating hallucination in LLMs. The dataset was constructed through a structured pipeline involving quality assurance, filtering for clarity and factual validity, and model-driven selection to retain questions that consistently trigger hallucinations. Each question is linked to verified ground-truth evidence, answer explanations, and multi-label annotations. Using the HalluScore benchmark, we conduct a comprehensive empirical analysis of hallucination patterns across 17 Arabic, multilingual, and reasoning LLMs. Moreover, we provide high-quality human annotations identifying hallucinated, non-hallucinated, and partially hallucinated responses of all evaluated LLMs. These results suggest that hallucination in Arabic LLMs extends beyond factual inaccuracies, encompassing challenges related to cultural understanding, linguistic reasoning, and logical consistency. We release HalluScore to support future research on improving the reliability and cultural competence of LLMs in Arabic.

幻觉评测阿拉伯语大模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。