Dr.V通过分层分析定位视频幻觉,提升大模型理解可靠性。
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- 构建感知-时序-认知三级框架,逐层诊断视频幻觉
- 在10,000个标注实例上验证,显著提升幻觉识别准确率
- 适合关注视频生成可信度与模型可解释性的研究者
大型视频模型(LVMs)在视频理解方面取得显著进展,但仍存在与输入视频内容冲突的幻觉问题。为解决该问题,我们提出Dr.V,一种涵盖感知、时序和认知层级的分层诊断框架,通过细粒度时空定位来识别视频幻觉。Dr.V包含两个核心组件:基准数据集Dr.V-Bench和卫星视频代理Dr.V-Agent。Dr.V-Bench包含从4,974段视频中抽取的10,000个实例,覆盖多样任务,并配有详细的时空标注。Dr.V-Agent通过在感知和时序层面系统应用细粒度时空定位,再进行认知层面推理,模拟人类视频理解过程,有效识别幻觉。大量实验表明,Dr.V-Agent在诊断幻觉的同时提升了可解释性与可靠性,为真实场景下的鲁棒视频理解提供了实用蓝图。所有数据与代码已公开于https://github.com/Eurekaleo/Dr.V。
原文摘要 · Abstract (English)
Recent advancements in large video models (LVMs) have significantly enhance video understanding. However, these models continue to suffer from hallucinations, producing content that conflicts with input videos. To address this issue, we propose Dr.V, a hierarchical framework covering perceptive, temporal, and cognitive levels to diagnose video hallucination by fine-grained spatial-temporal grounding. Dr.V comprises of two key components: a benchmark dataset Dr.V-Bench and a satellite video agent Dr.V-Agent. Dr.V-Bench includes 10k instances drawn from 4,974 videos spanning diverse tasks, each enriched with detailed spatial-temporal annotation. Dr.V-Agent detects hallucinations in LVMs by systematically applying fine-grained spatial-temporal grounding at the perceptive and temporal levels, followed by cognitive level reasoning. This step-by-step pipeline mirrors human-like video comprehension and effectively identifies hallucinations. Extensive experiments demonstrate that Dr.V-Agent is effective in diagnosing hallucination while enhancing interpretability and reliability, offering a practical blueprint for robust video understanding in real-world scenarios. All our data and code are available at https://github.com/Eurekaleo/Dr.V.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。