arXiv:2602.06920cs.CLcs.AI2026-02被引 3

构建多语言多任务幻觉检测基准,揭示模型在不同语言中的幻觉表现差异。

Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs

  • 设计跨四语言、双任务的幻觉数据集,区分实体、关系、句子级幻觉。
  • 中文与英文表现最佳,印地语检测准确率最低,对话摘要任务更易出错。
  • 适合研究多语言模型幻觉机制及评估工具开发的研究者使用。

大语言模型中的幻觉问题仍是持续挑战,尤其在多语言和生成场景下难以保持事实一致性。尽管近期模型在以英语为主的基准上表现良好,但其在不同语言、任务和幻觉类型下的行为仍不明确。本文提出 Halluverse-M^3,一个用于系统分析多语言、多任务、多类别幻觉的数据集。该数据集涵盖英语、阿拉伯语、印地语和土耳其语,支持问答与对话摘要两项生成任务,并明确区分实体级、关系级和句子级幻觉。幻觉输出通过受控编辑生成,并经人工标注验证,确保原始内容与幻觉输出间的清晰对应。我们在此数据集上评估了多种主流开源与专有模型的细粒度幻觉检测能力。结果表明:问答任务普遍比对话摘要更容易处理;即使最强模型也难以识别句子级幻觉。性能在英语中最高,低资源语言中下降,印地语检测准确率最低。整体而言,Halluverse-M^3为多语言多任务幻觉研究提供了真实且具有挑战性的基准。数据集已公开,供后续研究使用。

原文摘要 · Abstract (English)

Hallucinations in large language models remain a persistent challenge, particularly in multilingual and generative settings where factual consistency is difficult to maintain. While recent models show strong performance on English-centric benchmarks, their behavior across languages, tasks, and hallucination types is not yet well understood. In this work, we introduce Halluverse-M^3, a dataset designed to enable systematic analysis of hallucinations across multiple languages, multiple generation tasks, and multiple hallucination categories. Halluverse-M^3 covers four languages, English, Arabic, Hindi, and Turkish, and supports two generation tasks: question answering and dialogue summarization. The dataset explicitly distinguishes between entity-level, relation-level, and sentence-level hallucinations. Hallucinated outputs are constructed through a controlled editing process and validated by human annotators, ensuring clear alignment between original content and hallucinated generations. Using this dataset, we evaluate a diverse set of contemporary open-source and proprietary language models on fine-grained hallucination detection. Our results show that question answering is consistently easier than dialogue summarization, while sentence-level hallucinations remain challenging even for the strongest models. Performance is highest in English and degrades in lower-resource languages, with Hindi exhibiting the lowest detection accuracy. Overall, Halluverse-M^3 provides a realistic and challenging benchmark for studying hallucinations in multilingual, multi-task settings. We release the dataset to support future research on hallucination detection and mitigation\footnote{https://huggingface.co/datasets/sabdalja/HalluVerse-M3}.

幻觉检测多语言大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。