arXiv:2411.16771cs.CV2024-11被引 25

构建视频幻觉评估基准,揭示视觉大模型在时序理解中的严重幻觉问题。

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs

  • 基于视频时序特性设计多层级幻觉标注,实现细粒度幻觉评估
  • 提出排序任务让模型判断不同描述的幻觉程度,发现现有模型普遍误判
  • 适合关注视频理解、模型可靠性与幻觉检测的研究者使用

视觉大语言模型(VLLMs)被广泛认为容易产生幻觉。现有研究主要集中在图像输入,对视频幻觉的探索有限。当前评估方法难以捕捉生成响应中细微的错误,而这些错误常因视频丰富的时空动态而加剧。为此,我们提出VidHal,一个专门用于评估视频类幻觉的基准。VidHal通过跨多种常见时序特征自举视频实例构建,并精心设计了代表不同程度幻觉的字幕。为实现细粒度评估,我们提出一种新的字幕排序任务,要求模型按幻觉程度对字幕进行排序。我们在VidHal上进行了大规模实验,全面评估了多种模型。结果揭示了现有VLLMs在生成幻觉方面的显著缺陷。本基准旨在推动对VLLM能力的全面理解,特别是幻觉问题,并促进先进VLLMs的发展以缓解该问题。

原文摘要 · Abstract (English)

Vision Large Language Models (VLLMs) are widely acknowledged to be prone to hallucinations. Existing research addressing this problem has primarily been confined to image inputs, with limited exploration of video-based hallucinations. Furthermore, current evaluation methods fail to capture nuanced errors in generated responses, which are often exacerbated by the rich spatiotemporal dynamics of videos. To address this, we introduce VidHal, a benchmark specially designed to evaluate video-based hallucinations in VLLMs. VidHal is constructed by bootstrapping video instances across a wide range of common temporal aspects. A defining feature of our benchmark lies in the careful creation of captions which represent varying levels of hallucination associated with each video. To enable fine-grained evaluation, we propose a novel caption ordering task requiring VLLMs to rank captions by hallucinatory extent. We conduct extensive experiments on VidHal and comprehensively evaluate a broad selection of models. Our results uncover significant limitations in existing VLLMs regarding hallucination generation. Through our benchmark, we aim to inspire further research on 1) holistic understanding of VLLM capabilities, particularly regarding hallucination, and 2) extensive development of advanced VLLMs to alleviate this problem.

视频理解幻觉检测大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。