arXiv:2503.01670cs.CLcs.AI2025-03ACL被引 9

检验大模型在混合上下文幻觉检测中的表现,发现其判断易受自身知识偏见影响。

Evaluating LLMs' Assessment of Mixed-Context Hallucination Through the Lens of Summarization

  • 以摘要生成任务为场景,评估大模型对混合上下文幻觉的识别能力。
  • 大模型更难识别事实性幻觉,性能受自身知识干扰严重。
  • 适合关注大模型评测可靠性与幻觉检测机制的研究者阅读。

随着大语言模型(LLMs)的快速发展,将大模型作为评判者(LLM-as-a-judge)已成为文本质量评估的常用方法,包括幻觉检测。以往研究多集中于单上下文评估(如论述一致性或世界事实性),但现实中的幻觉通常涉及混合上下文,相关评估仍不充分。本文以摘要生成为代表性任务,全面评估大模型在区分事实性与非事实性幻觉方面的混合上下文幻觉检测能力。通过在不同规模的直接生成与基于检索的模型上进行大量实验,主要发现:(1)大模型内在知识引入了固有偏见,影响幻觉评估;(2)这些偏见尤其损害事实性幻觉的检测,形成显著性能瓶颈;(3)根本挑战在于如何有效利用知识,平衡模型内生知识与外部上下文,实现准确的混合上下文幻觉评估。

原文摘要 · Abstract (English)

With the rapid development of large language models (LLMs), LLM-as-a-judge has emerged as a widely adopted approach for text quality evaluation, including hallucination evaluation. While previous studies have focused exclusively on single-context evaluation (e.g., discourse faithfulness or world factuality), real-world hallucinations typically involve mixed contexts, which remains inadequately evaluated. In this study, we use summarization as a representative task to comprehensively evaluate LLMs' capability in detecting mixed-context hallucinations, specifically distinguishing between factual and non-factual hallucinations. Through extensive experiments across direct generation and retrieval-based models of varying scales, our main observations are: (1) LLMs' intrinsic knowledge introduces inherent biases in hallucination evaluation; (2) These biases particularly impact the detection of factual hallucinations, yielding a significant performance bottleneck; (3) The fundamental challenge lies in effective knowledge utilization, balancing between LLMs' intrinsic knowledge and external context for accurate mixed-context hallucination evaluation.

幻觉检测大模型评测摘要生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。