arXiv:2508.09584cs.CV2025-08被引 4

构建可扩展的细粒度幻觉评估基准,精准检测视觉语言模型幻觉问题。

SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs

论文配图:SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
图 1 · 摘自论文原文
  • 自动化生成数据,支持可控且多样化的评估场景。
  • 覆盖30多千图像指令对,涵盖12类感知与6个知识领域。
  • 适合关注模型可靠性与鲁棒性的研究人员使用。

尽管大视觉语言模型(LVLMs)发展迅速,但仍存在幻觉问题,即生成内容与输入或已知世界知识不一致,分别对应忠实性与事实性幻觉。以往研究多在粗粒度层面(如物体级)评估忠实性幻觉,缺乏细粒度分析。现有基准常依赖高成本人工标注或重复使用公开数据集,存在可扩展性与数据泄露风险。为此,我们提出自动化数据构建流程,实现可扩展、可控、多样化的评估数据生成;设计分层幻觉诱导框架,通过输入扰动模拟真实噪声场景。结合上述设计,构建了SHALE——一个可扩展的幻觉评估基准,通过细粒度分类体系评估忠实性与事实性幻觉。SHALE包含超过3万张图像-指令对,覆盖12个代表性视觉感知维度(忠实性)和6个知识领域(事实性),兼顾清洁与噪声场景。在20多个主流LVLM上的实验表明,存在显著的事实性幻觉,且对语义扰动高度敏感。

原文摘要 · Abstract (English)

Despite rapid advances, Large Vision-Language Models (LVLMs) still suffer from hallucinations, i.e., generating content inconsistent with input or established world knowledge, which correspond to faithfulness and factuality hallucinations, respectively. Prior studies primarily evaluate faithfulness hallucination at a rather coarse level (e.g., object-level) and lack fine-grained analysis. Additionally, existing benchmarks often rely on costly manual curation or reused public datasets, raising concerns about scalability and data leakage. To address these limitations, we propose an automated data construction pipeline that produces scalable, controllable, and diverse evaluation data. We also design a hierarchical hallucination induction framework with input perturbations to simulate realistic noisy scenarios. Integrating these designs, we construct SHALE, a Scalable HALlucination Evaluation benchmark designed to assess both faithfulness and factuality hallucinations via a fine-grained hallucination categorization scheme. SHALE comprises over 30K image-instruction pairs spanning 12 representative visual perception aspects for faithfulness and 6 knowledge domains for factuality, considering both clean and noisy scenarios. Extensive experiments on over 20 mainstream LVLMs reveal significant factuality hallucinations and high sensitivity to semantic perturbations.

幻觉评估视觉语言模型可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。