arXiv:2505.08455cs.CV2025-05被引 3

评测大模型视频因果推理能力,发现其难以处理长序列因果关系。

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models

  • 用打乱步骤的日常活动视频构建新基准,测试模型事件排序能力。
  • 现有大模型在该任务上准确率低,最高提升25.2%(引入新方法后)。
  • 模型主要依赖语言知识而非视觉推理,适合研究视频理解与因果建模者。

尽管视频理解取得进展,大型视频语言模型(LVLMs)在基于视频的因果推理能力仍缺乏系统评估,主要因缺乏专门的基准。为此,我们提出新基准VCRBench,使用简单日常活动的程序化视频,将步骤打乱并每段捕捉关键因果事件,测试模型是否能识别、推理并正确排序完成特定目标所需事件。该基准避免语言捷径(如多选题),也规避开放式问答的评估难题。对主流LVLM的评估显示,模型在长序列因果推理上表现不佳,主要因难以从视觉中建模长程因果依赖。为此,我们提出识别-推理分解(RRD)框架,将任务拆分为视频识别与因果推理两阶段。实验表明,RRD在VCRBench上显著提升准确率,最高达25.2%。深入分析揭示,模型在复杂任务中主要依赖语言知识而非视觉推理。

原文摘要 · Abstract (English)

Despite recent advances in video understanding, the capabilities of Large Video Language Models (LVLMs) to perform video-based causal reasoning remains underexplored, largely due to the absence of relevant and dedicated benchmarks for evaluating causal reasoning in visually grounded and goal-driven settings. To fill this gap, we introduce a novel benchmark named Video-based long-form Causal Reasoning (VCRBench). We create VCRBench using procedural videos of simple everyday activities, where the steps are deliberately shuffled with each clip capturing a key causal event, to test whether LVLMs can identify, reason about, and correctly sequence the events needed to accomplish a specific goal. Moreover, the benchmark is carefully designed to prevent LVLMs from exploiting linguistic shortcuts, as seen in multiple-choice or binary QA formats, while also avoiding the challenges associated with evaluating open-ended QA. Our evaluation of state-of-the-art LVLMs on VCRBench suggests that these models struggle with video-based long-form causal reasoning, primarily due to their difficulty in modeling long-range causal dependencies directly from visual observations. As a simple step toward enabling such capabilities, we propose Recognition-Reasoning Decomposition (RRD), a modular approach that breaks video-based causal reasoning into two sub-tasks of video recognition and causal reasoning. Our experiments on VCRBench show that RRD significantly boosts accuracy on VCRBench, with gains of up to 25.2%. Finally, our thorough analysis reveals interesting insights, for instance, that LVLMs primarily rely on language knowledge for complex video-based long-form causal reasoning tasks.

视频理解因果推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。