arXiv:2608.03782cs.AI2026-08

构建多模态幻觉评估新基准,系统评测模型在知识层面的错误。

KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation

论文配图:KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation
图 1 · 摘自论文原文
  • 设计配对正负问题,控制比较感知错误与知识错误。
  • 14个模型在知识维度表现最差,负向问题下性能普遍下降。
  • 适合研究多模态大模型可信性与评估框架的学者使用。

幻觉仍是构建可信多模态大语言模型(MLLMs)的关键挑战。现有基准主要关注实体、属性和关系幻觉,而知识相关错误常被单独研究,缺乏跨维度统一评估框架。为此,我们提出 extbf{KnowHal},将知识幻觉纳入多模态幻觉评估,覆盖实体、属性、关系和知识四个维度。KnowHal在共享图像和实体上构建配对正负问题,实现对感知错误、知识错误及虚假前提接受的可控对比。该基准包含1,800个样本,覆盖10个领域和50个类别,通过半自动化流程(LLM辅助、CLIP过滤、人工验证)构建。我们在14个代表性MLLM上评估KnowHal并进行深入分析。结果表明,知识维度对几乎所有模型构成最大挑战,多数模型在负向问题上性能显著下降,揭示其对虚假前提鲁棒性不足。通过四维统一与配对设计,KnowHal填补了现有评估框架的重要空白,实现对MLLM幻觉更全面的评估。

原文摘要 · Abstract (English)

Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks mainly focus on entity, attribute, and relation hallucinations, knowledge-related failures are often investigated separately, lacking a unified evaluation framework across different hallucination dimensions. To overcome this, we propose \textbf{KnowHal}, a benchmark that explicitly incorporates knowledge hallucination into multimodal hallucination evaluation spanning four dimensions: entity, attribute, relation, and knowledge. KnowHal constructs paired positive and negative questions over shared images and entities, enabling controlled comparisons among perceptual errors, knowledge-related errors, and false-premise acceptance. The benchmark contains 1,800 samples across 10 domains and 50 categories, constructed through a semi-automated pipeline combining LLM assistance, CLIP-based filtering, and human verification. We evaluate 14 representative MLLMs on KnowHal and conduct extensive analyses. Results show that the knowledge dimension consistently presents the greatest challenge for nearly all evaluated models, while most models exhibit substantial performance degradation on negative questions, revealing limited robustness to false premises. By unifying four hallucination dimensions with paired question design, KnowHal addresses an important gap in existing evaluation frameworks and enables a more comprehensive assessment of hallucinations in MLLMs.

多模态幻觉评估知识幻觉基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。