构建首个区分内外部幻觉的通用大模型幻觉评测基准
HalluLens: LLM Hallucination Benchmark

- 按内外部幻觉分类,建立统一评测框架
- 设计可动态重生成的数据集,防止数据泄露导致评估失真
- 揭示现有评测局限性,推动可信AI研究
大语言模型常生成与用户输入或训练数据不符的内容,即‘幻觉’,这削弱了用户信任并阻碍生成式AI应用。本文提出一个全面的幻觉评测基准,包含新提出的外部幻觉任务与现有内在幻觉任务,基于清晰的幻觉分类体系。评测的核心挑战在于缺乏统一框架,因定义和分类不一致。本文将幻觉从‘事实性’中分离,提出内外部幻觉的明确分类,促进研究一致性。外部幻觉(生成内容与训练数据不一致)随模型演进而日益重要。基准采用动态测试集生成机制,防止数据泄露并确保评估鲁棒性。同时分析现有基准的局限性与饱和现象。工作目标为:(1) 建立幻觉清晰分类体系;(2) 引入新外部幻觉任务,数据可动态再生以避免泄漏;(3) 综合分析现有基准,区分其与事实性评估的本质差异。
原文摘要 · Abstract (English)
Large language models (LLMs) often generate responses that deviate from user input or training data, a phenomenon known as "hallucination." These hallucinations undermine user trust and hinder the adoption of generative AI systems. Addressing hallucinations is essential for the advancement of LLMs. This paper introduces a comprehensive hallucination benchmark, incorporating both new extrinsic and existing intrinsic evaluation tasks, built upon clear taxonomy of hallucination. A major challenge in benchmarking hallucinations is the lack of a unified framework due to inconsistent definitions and categorizations. We disentangle LLM hallucination from "factuality," proposing a clear taxonomy that distinguishes between extrinsic and intrinsic hallucinations, to promote consistency and facilitate research. Extrinsic hallucinations, where the generated content is not consistent with the training data, are increasingly important as LLMs evolve. Our benchmark includes dynamic test set generation to mitigate data leakage and ensure robustness against such leakage. We also analyze existing benchmarks, highlighting their limitations and saturation. The work aims to: (1) establish a clear taxonomy of hallucinations, (2) introduce new extrinsic hallucination tasks, with data that can be dynamically regenerated to prevent saturation by leakage, (3) provide a comprehensive analysis of existing benchmarks, distinguishing them from factuality evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。