arXiv:2602.19643cs.CL2026-02Conference of the …

用知识图谱构建动态问答,更全面评估大模型幻觉问题。

KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge

  • 基于知识图谱生成多维度动态问题,提升评测覆盖范围。
  • 25个前沿模型测试显示,模型规模影响幻觉类型与程度。
  • 可自动检测回答偏差,适合研究幻觉机制与模型评估者使用。

大型语言模型(LLMs)能生成流畅自然的语言,但其内容常包含细微的幻觉。现有基准测试受限于静态、单一的问题,导致评估覆盖面窄且易产生误导。我们提出KGHaluaBench,一个基于知识图谱的幻觉评测基准,可从广度和深度两方面评估LLM的知识真实性。该框架利用知识图谱动态生成复杂多面的问题,并通过统计方法估计难度以缓解流行性偏差。自动化验证流程在概念和正确性层面检测模型回答中的回避行为与错误,识别不同类型的幻觉。我们评估了25个前沿模型,引入新的准确率与幻觉度量指标,揭示了不同模型规模下导致幻觉的关键知识因素。KGHaluaBench已开源,支持未来幻觉缓解研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) possess a remarkable capacity to generate persuasive and intelligible language. However, coherence does not equate to truthfulness, as the responses often contain subtle hallucinations. Existing benchmarks are limited by static and narrow questions, leading to limited coverage and misleading evaluations. We present KGHaluBench, a Knowledge Graph-based hallucination benchmark that assesses LLMs across the breadth and depth of their knowledge, providing a fairer and more comprehensive insight into LLM truthfulness. Our framework utilises the KG to dynamically construct challenging, multifaceted questions, whose difficulty is then statistically estimated to address popularity bias. Our automated verification pipeline detects abstentions and verifies the LLM's response at both conceptual and correctness levels to identify different types of hallucinations. We evaluate 25 frontier models, using novel accuracy and hallucination metrics. The results provide a more interpretable insight into the knowledge factors that cause hallucinations across different model sizes. KGHaluBench is publicly available to support future developments in hallucination mitigation.

幻觉评测知识图谱大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。