arXiv:2503.18923cs.CV2025-03被引 11

首个专用于视频事实性评估的基准,检验大模型能否准确理解视频中的多跳事实。

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

  • 设计需融合外部知识的多跳问题,要求逐步推理并定位视频时间片段。
  • 顶尖模型事实准确率仅66.3%,多数模型过度自信,推理能力严重不足。
  • 适合关注视频理解真实性、模型可解释性的研究者与开发者使用。

大型视频语言模型(LVLMs)在多模态理解方面取得进展,但其对视频内容的事实性支撑仍是一大挑战。为此,我们提出Video SimpleQA,首个面向视频场景的事实性评估综合基准。该基准具有四大特点:1)需整合视频外的外部知识;2)采用多跳事实型问题,每题包含多个明确事实,要求严格基于事实而非假设或主观推断,并附带逐跳子问题以实现细粒度评估;3)答案为简明无歧义的确定性回答,降低评分波动;4)要求答案基于视频中的一个或多个时间片段,而非单帧。我们对33个前沿LVLMs进行了全面评测,发现:1)当前模型事实性表现不佳,最优模型o3仅达66.3% F-score;2)多数模型生成时过度自信,自评置信度远超实际准确率;3)检索增强生成虽提升性能,但增加推理延迟;4)多跳问答性能显著低于单跳子问题,首跳的对象/事件识别成为主要瓶颈。Video SimpleQA旨在成为视频事实性评估的基石,推动模型向真实世界可验证的语义理解发展。

原文摘要 · Abstract (English)

Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains a critical unsolved challenge. To address this gap, we introduce Video SimpleQA, the first comprehensive benchmark tailored for factuality evaluation in video contexts. Our work differs from existing video benchmarks through the following key features: 1) Knowledge required: demanding integration of external knowledge beyond the video's explicit narrative; 2) Multi-hop fact-seeking question: Each question involves multiple explicit facts and requires strict factual grounding without hypothetical or subjective inferences. We also include per-hop single-fact-based sub-QAs alongside final QAs to enable fine-grained, stepby-step evaluation; 3) Short-form definitive answer: Answers are crafted as unambiguous and definitively correct in a short format with minimal scoring variance; 4) Temporal grounded required: Requiring answers to rely on one or more temporal segments in videos, rather than single frames. We extensively evaluate 33 state-of-the-art LVLMs and summarize key findings as follows: 1) Current LVLMs exhibit notable deficiencies in factual adherence, with the best-performing model o3 merely achieving an F-score of 66.3%; 2) Most LVLMs are overconfident in what they generate, with self-stated confidence exceeding actual accuracy; 3) Retrieval-augmented generation demonstrates consistent improvements at the cost of additional inference time overhead; 4) Multi-hop QA demonstrates substantially degraded performance compared to single-hop sub-QAs, with first-hop object or event recognition emerging as the primary bottleneck. We position Video SimpleQA as the cornerstone benchmark for video factuality assessment, aiming to steer LVLM development toward verifiable grounding in real-world contexts.

视频理解事实性评估多跳推理大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。