构建多语言细粒度幻觉数据集,助力评估大模型事实性
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
- 用大模型注入幻觉后经人工校验,构建英阿土三语数据集
- 首次系统标注实体、关系、句子级幻觉,覆盖多种语言
- 适合研究模型幻觉检测或跨语言事实性评估的学者
大语言模型在各类场景中应用日益广泛,但仍易生成非事实内容,即“幻觉”。现有幻觉数据集通常无法捕捉多语言环境下的细粒度幻觉。本文提出 HalluVerse25,一个涵盖英语、阿拉伯语和土耳其语的多语言幻觉数据集,对实体级、关系级和句子级幻觉进行细粒度标注。数据构建采用大模型注入幻觉至真实传记句,再经严格人工标注以保证质量。我们在该数据集上评估多个大语言模型,揭示了商业模型在不同语境下检测幻觉的表现差异。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used in various contexts, yet remain prone to generating non-factual content, commonly referred to as "hallucinations". The literature categorizes hallucinations into several types, including entity-level, relation-level, and sentence-level hallucinations. However, existing hallucination datasets often fail to capture fine-grained hallucinations in multilingual settings. In this work, we introduce HalluVerse25, a multilingual LLM hallucination dataset that categorizes fine-grained hallucinations in English, Arabic, and Turkish. Our dataset construction pipeline uses an LLM to inject hallucinations into factual biographical sentences, followed by a rigorous human annotation process to ensure data quality. We evaluate several LLMs on HalluVerse25, providing valuable insights into how proprietary models perform in detecting LLM-generated hallucinations across different contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。