arXiv:2506.10857cs.CVcs.AI2025-06ICCV被引 27

首个评估长视频多步推理能力的基准,聚焦时间逻辑与过程有效性。

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

  • 构建960段平均时长1.6小时的长视频,含8243个问答对与25106条带时间戳的推理步骤。
  • 提出多阶段评估流程,用基于LLM的进度评分指标衡量推理链质量。
  • 适合研究视频理解、多步推理和长时序建模的学者与工程师使用。

我们提出VRBench,首个专为评估大模型在长叙事视频中多步推理能力而设计的基准,弥补了现有评估忽略时间推理与过程有效性的不足。该基准包含960段长视频(平均时长1.6小时),以及8,243个人工标注的多步问答对和25,106条带时间戳的推理步骤。视频通过多阶段筛选流程(包括专家间一致性评审)精心选取,以确保情节连贯性。我们开发了一种人机协同框架,生成需多个时间锚定步骤的连贯推理链,涵盖七类任务(如事件归因、隐含推断)。VRBench设计了多阶段评估流程,在结果层(如多选题)之外,引入基于LLM的进度级评分指标,从多个维度综合评估推理链质量。通过对12个LLM和19个VLM在VRBench上的广泛评测,我们开展了深入分析,为多步推理领域提供了宝贵洞见。

原文摘要 · Abstract (English)

We present VRBench, the first long narrative video benchmark crafted for evaluating large models' multi-step reasoning capabilities, addressing limitations in existing evaluations that overlook temporal reasoning and procedural validity. It comprises 960 long videos (with an average duration of 1.6 hours), along with 8,243 human-labeled multi-step question-answering pairs and 25,106 reasoning steps with timestamps. These videos are curated via a multi-stage filtering process including expert inter-rater reviewing to prioritize plot coherence. We develop a human-AI collaborative framework that generates coherent reasoning chains, each requiring multiple temporally grounded steps, spanning seven types (e.g., event attribution, implicit inference). VRBench designs a multi-phase evaluation pipeline that assesses models at both the outcome and process levels. Apart from the MCQs for the final results, we propose a progress-level LLM-guided scoring metric to evaluate the quality of the reasoning chain from multiple dimensions comprehensively. Through extensive evaluations of 12 LLMs and 19 VLMs on VRBench, we undertake a thorough analysis and provide valuable insights that advance the field of multi-step reasoning.

视频理解多步推理长视频基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。