arXiv:2607.20219cs.CL2026-07被引 2

首个细粒度阿拉伯语幻觉检测基准,支持定位与解释错误。

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

论文配图:HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering
图 1 · 摘自论文原文
  • 构建2400个专家标注的阿拉伯语问答对,含错误片段与解释。
  • 模型在错误定位上表现最差(F1-Sp仅0.516),验证能力较强。
  • 适合研究阿拉伯语幻觉检测、可解释AI的学者和开发者。

大语言模型能生成流畅的阿拉伯语答案,但事实性错误仍难以检测、定位、解释和验证。现有幻觉评估基准多仅提供回答级标签,缺乏对具体错误内容的识别、错误原因说明或正确答案选择支持。我们提出HalluTruthQA,一个面向阿拉伯语问答的细粒度幻觉评估基准。该基准包含2400个跨四大知识领域(伊斯兰知识、历史、科学、地理)的专家标注示例,每个样本包含阿拉伯语问题、模型生成答案、已验证参考答案、二元幻觉标签及六个候选答案用于事实验证。幻觉答案额外标注字符级错误片段、人工撰写解释,并区分宏观与微观幻觉类型。我们在零样本设置下评估四个开源LLM(ALLaM-7B、Falcon-H1R-7B、Qwen3-32B、SILMA)在幻觉检测、片段级定位、事实验证和解释评估四项任务上的表现。结果显示各任务反映不同能力:无模型在所有任务中均最优。最佳成绩为:检测0.880宏平均F1,定位0.516 F1-Sp,事实验证0.852 LO-Score,解释评估0.644。结果表明,幻觉评估应从回答级检测转向错误定位、验证与解释。

原文摘要 · Abstract (English)

Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer. We introduce HalluTruthQA, a fine-grained benchmark for hallucination evaluation in Arabic question answering. The benchmark contains 2,400 expert-curated examples across four knowledge-intensive domains: Islamic knowledge, history, science, and geography. Each example pairs an Arabic question and a model-generated answer with a verified reference answer, a binary hallucination label, and six candidate answers for factual verification. Hallucinated answers additionally include character-level erroneous spans, human-written explanations, and macro- and micro-level hallucination types. We evaluate four open-source LLMs, ALLaM-7B, Falcon-H1R-7B, Qwen3-32B, and SILMA, in a zero-shot setting across hallucination detection, span-level localization, factual verification, and explanation evaluation. Results show that these tasks capture different abilities: no single model performs best across all tasks. The best scores are 0.880 Macro-F1 for detection, 0.516 F1-Sp for localization, 0.852 LO-Score for factual verification, and 0.644 for explanation evaluation. These findings show that hallucination evaluation should move beyond response-level detection toward the localization, verification, and explanation of factual errors.

幻觉检测阿拉伯语细粒度评估可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。