构建4000条阿拉伯语幻觉检测数据集,支持细粒度错误定位与验证。
HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification

- 专家标注4000个问答对,含错误片段与原因说明
- 1643条含幻觉,覆盖伊斯兰、历史、科学、地理四领域
- 适合评估阿拉伯语模型事实可靠性与错误定位能力
大型语言模型生成流畅的阿拉伯语回答时常引入难以识别的事实性错误。现有阿拉伯语幻觉资源多仅对整个回答给出二元标签,缺乏对具体错误内容、错误原因及正确答案的详细信息。本文提出HalluTruthQA-4K,是原资源的扩展版本,包含4,000个跨伊斯兰知识、历史、科学和地理四个知识密集型领域的专家标注阿拉伯语问答实例。每条数据包含一个阿拉伯语问题、模型生成的回答、经验证的参考答案以及五个合理干扰项。幻觉回答额外标注字符级错误片段、人工编写的解释及层级化幻觉类型。该语料库包含1,643条幻觉响应与2,357条非幻觉响应,共标注1,843个错误片段。本文详述资源构建与标注方法,包括问题筛选、受控回答生成、候选构造、专家标注、独立验证、争议解决与质量控制流程。同时记录标注指南、分类体系、数据格式、标注者一致性及语料统计。HalluTruthQA-4K可复用于幻觉检测、错误片段定位、解释生成、事实验证及阿拉伯语模型事实可靠性的综合评估。
原文摘要 · Abstract (English)
Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hallucination resources often assign a binary label to an entire response, indicating whether it is hallucinated or non-hallucinated, but provide limited information about the exact erroneous content, the reason for the error, or the correct factual answer. We present HalluTruthQA-4K, an expanded version of the HalluTruthQA resource containing 4,000 expert-curated Arabic question-answering instances across four knowledge-intensive domains: Islamic knowledge, history, science, and geography. Serving as the official dataset for Track 2 of the HalluScoring 2026 shared task, HalluTruthQA-4K extends our original corpus to 4,000 instances. Each instance pairs an Arabic question with a model-generated response, a verified reference answer, and five plausible distractors. Hallucinated responses are additionally annotated with character-level erroneous spans, human-written explanations, and hierarchical hallucination types. The corpus contains 1,643 hallucinated and 2,357 non-hallucinated responses, with 1,843 annotated erroneous spans. We describe the resource construction and annotation methodology, including question selection, controlled answer generation, candidate construction, expert annotation, independent verification, adjudication, and quality control. We also document the annotation guidelines, taxonomy, data format, inter-annotator agreement, and corpus statistics. HalluTruthQA-4K provides a reusable resource for hallucination detection, span-level error localization, explanation generation, factual verification, and the broader evaluation of factual reliability in Arabic language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。