arXiv:2504.07956cs.CVcs.AI2025-04被引 37

构建视频链式推理评估基准,揭示大模型在视频理解中的真实短板

VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

  • 设计带步骤标注的视频推理问答对,区分感知与推理环节
  • 顶级模型o1仅达62.8%推理得分,多数模型低于40%
  • 发现模型在时序空间信息处理上存在明显瓶颈,适合研究视频推理的学者

链式思维(CoT)推理显著提升了大语言模型和大视觉语言模型的能力,但针对视频场景的严谨评估框架仍属空白。现有视频基准难以准确评估推理过程,无法判断失败源于感知还是推理能力不足。为此,我们提出VCR-Bench,一个全面评估大视觉语言模型视频链式思维推理能力的新基准。该基准包含859段涵盖多种内容与时长的视频,以及1,034个高质量问题-答案对。每个答案对均经人工标注逐步链式推理理由,并标记每一步所属的感知或推理能力类型。我们设计了七个任务维度,并提出基于分步标注的链式思维得分(CoT score)来评估完整推理流程。在VCR-Bench上的大量实验表明当前模型存在显著局限:即使表现最优的o1模型也仅达到62.8%的CoT得分和56.7%的准确率,多数模型得分低于40%。实验显示,模型在感知类步骤得分普遍低于推理类步骤,暴露了大模型在复杂视频时序-空间信息处理上的关键瓶颈。链式思维得分与准确率之间存在强正相关性,验证了评估框架的有效性,并强调链式推理在解决复杂视频任务中的核心作用。我们希望VCR-Bench能成为标准化评估工具,揭示复杂视频推理任务中的实际缺陷。

原文摘要 · Abstract (English)

The advancement of Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs) and large vision-language models (LVLMs). However, a rigorous evaluation framework for video CoT reasoning remains absent. Current video benchmarks fail to adequately assess the reasoning process and expose whether failures stem from deficiencies in perception or reasoning capabilities. Therefore, we introduce VCR-Bench, a novel benchmark designed to comprehensively evaluate LVLMs' Video Chain-of-Thought Reasoning capabilities. VCR-Bench comprises 859 videos spanning a variety of video content and durations, along with 1,034 high-quality question-answer pairs. Each pair is manually annotated with a stepwise CoT rationale, where every step is tagged to indicate its association with the perception or reasoning capabilities. Furthermore, we design seven distinct task dimensions and propose the CoT score to assess the entire CoT process based on the stepwise tagged CoT rationals. Extensive experiments on VCR-Bench highlight substantial limitations in current LVLMs. Even the top-performing model, o1, only achieves a 62.8% CoT score and an 56.7% accuracy, while most models score below 40%. Experiments show most models score lower on perception than reasoning steps, revealing LVLMs' key bottleneck in temporal-spatial information processing for complex video reasoning. A robust positive correlation between the CoT score and accuracy confirms the validity of our evaluation framework and underscores the critical role of CoT reasoning in solving complex video reasoning tasks. We hope VCR-Bench to serve as a standardized evaluation framework and expose the actual drawbacks in complex video reasoning task.

视频推理链式思维评估基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。