arXiv:2604.05623cs.CVcs.CL2026-04被引 2

构建首个针对长图像描述中密集幻觉定位的基准测试

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions

  • 设计跨五领域1000张图像的细粒度标注数据集
  • 平均200+词长描述,支持逐词级幻觉定位评估
  • 适合评估多模态大模型在长文本生成中的可靠性

准确检测并定位幻觉是保障图像描述高可靠性的关键任务。随着多模态大语言模型(MLLMs)的发展,图像描述已从简短语句演变为涵盖数百词的完整叙述。这一转变使挑战指数级增加:模型需在长篇上下文中精确定位错误片段或词语,而非仅识别整体不一致。然而,现有基准缺乏足够细粒度与领域多样性来评估此能力。为此,我们提出DetailVerifyBench,一个包含1000张高质量图像、覆盖五个不同领域的严格基准。其平均描述长度超过200词,且具备多种幻觉类型的逐标记标注,是当前长图像描述幻觉精准定位领域最具有挑战性的基准。数据集已公开于 https://zyx-hhnkh.github.io/DetailVerifyBench/。

原文摘要 · Abstract (English)

Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), captions have evolved from brief sentences into comprehensive narratives, often spanning hundreds of words. This shift exponentially increases the challenge: models must now pinpoint specific erroneous spans or words within extensive contexts, rather than merely flag response-level inconsistencies. However, existing benchmarks lack the fine granularity and domain diversity required to evaluate this capability. To bridge this gap, we introduce DetailVerifyBench, a rigorous benchmark comprising 1,000 high-quality images across five distinct domains. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as the most challenging benchmark for precise hallucination localization in the field of long image captioning to date. Our benchmark is available at https://zyx-hhnkh.github.io/DetailVerifyBench/.

幻觉检测图像描述多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。