arXiv:2606.24797cs.CVcs.AI2026-06

构建首个需精准定位时间证据的视频问答基准,检验模型是否真懂视频。

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

论文配图:EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
图 1 · 摘自论文原文
  • 提出EG-VQA基准,每题标注具体时间证据,要求答案与视频片段对齐。
  • 现有强模型在证据定位上表现差,说明答对≠真理解,差距明显。
  • 新模型EG-Reasoner通过显式监督提升证据推理,尤其擅长反事实问题。

视频大语言模型在视频问答(VideoQA)上取得进展,但现有评估多依赖答案正确性,忽视预测结果是否基于真实视频证据。为弥合这一断层,本文构建了证据锚定视频问答基准(EG-VQA),包含2,067个视频和11,838个问答对,每对均标注细粒度的时间证据,要求联合推理与精确证据定位。为此提出证据锚定F1(EG-F1)指标,同时衡量时间对齐与语义一致性。实验表明,即使强大专有模型也难以准确定位证据,暴露出答案正确性与证据可信度之间的根本差异。为此提出EG-Reasoner模型,通过显式监督进行证据推理,在开源模型中达到领先水平,尤其在反事实等推理密集型任务上优势显著。结果表明,仅靠规模扩展不足以实现可靠视频理解,结构化证据监督是构建可解释、可信视频问答系统的关键。

原文摘要 · Abstract (English)

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, existing benchmarks are predominantly evaluated through answer correctness, while the grounding of predictions in relevant video evidence remains largely unexamined. This disconnect between answer generation and evidence understanding motivates the construction of the Evidence-Grounded Video Question Answering Benchmark (EG-VQA), an open-ended evaluation protocol in which each QA pair is explicitly annotated with supporting temporal evidence, thereby requiring joint reasoning and precise evidence localization. EG-VQA is comprised of 2,067 videos and 11,838 QA pairs with fine-grained evidence annotations. To evaluate predicted evidence, Evidence-Grounded F1 (EG-F1) is introduced as a unified metric in which temporal alignment and semantic consistency against ground-truth evidence are jointly measured. Experimental evaluation reveals that even strong proprietary models struggle to accurately ground their predictions, exposing a fundamental discrepancy between answer correctness and faithful evidence localization. To bridge this gap, EG-Reasoner, an evidence-grounded reasoning model trained with explicit supervision, is proposed. State-of-the-art performance is achieved among open-source models, with results competitive against proprietary systems, particularly pronounced gains are observed on reasoning-intensive tasks such as counterfactual questions. These findings demonstrate that scaling alone is insufficient for robust video understanding and that structured evidence supervision is essential for the development of more reliable and interpretable VideoQA systems.

视频问答证据定位大模型评估可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。