arXiv:2508.19542cs.CV2025-08被引 13

首个评估多视频协同推理能力的基准,揭示当前模型在跨视频因果推理上的显著短板。

CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning

  • 构建三层次多视频关联任务,涵盖物体、事件与复杂推理
  • 顶尖模型在因果推理上仅达63.5%准确率,远低于人类91.3%表现
  • 揭示模型跨视频上下文保留差、实体歧义处理弱的核心瓶颈

尽管多模态大语言模型(MLLMs)在单视频任务(如视频问答)中表现优异,其在多视频间时空模式推理方面的能力仍存在关键缺口。该能力对多摄像头监控、跨视频流程学习等现实应用至关重要。为此,我们提出CVBench,首个专门用于严格评估跨视频关系推理的诊断性基准。CVBench包含1,000个问答对,涵盖三个层级:跨视频对象关联(识别共享实体)、跨视频事件关联(链接时间或因果事件链)、跨视频复杂推理(整合常识与领域知识)。数据源自五个领域多样化的视频集群(如体育、生活记录),要求模型分析并整合动态视觉流中的时空模式。对10+主流MLLM(包括GPT-4o、Gemini-2.0-flash、Qwen2.5-VL)在零样本或思维链提示下的广泛评估显示:即使顶级模型如GPT-4o,在因果推理任务上也仅达63.5%准确率,远低于人类91.3%的表现。分析揭示当前架构的根本缺陷,包括跨视频上下文保留不足与重叠实体辨析能力差。CVBench为多视频场景下的模式识别方法发展提供严谨框架,并为下一代模型设计提供架构启示。数据与评估代码已开源:https://github.com/Hokhim2/CVBench。

原文摘要 · Abstract (English)

While multimodal large language models (MLLMs) exhibit strong performance on single-video tasks (e.g., video question answering), their capability for spatiotemporal pattern reasoning across multiple videos remains a critical gap in pattern recognition research. However, this capability is essential for real-world applications, including multi-camera surveillance and cross-video procedural learning. To bridge this gap, we present CVBench, the first diagnostic benchmark designed to assess cross-video relational reasoning rigorously. CVBench comprises 1,000 question-answer pairs spanning three hierarchical tiers: cross-video object association (identifying shared entities), cross-video event association (linking temporal or causal event chains), and cross-video complex reasoning (integrating commonsense and domain knowledge). Built from five domain-diverse video clusters (e.g., sports, life records), the benchmark challenges models to analyze and integrate spatiotemporal patterns from dynamic visual streams. Extensive evaluation of 10+ leading MLLMs (including GPT-4o, Gemini-2.0-flash, Qwen2.5-VL) under zero-shot or chain-of-thought prompting paradigms. Key findings reveal stark performance gaps: even top models, such as GPT-4o, achieve only 63.5% accuracy on causal reasoning tasks, compared to the 91.3% accuracy of human performance. Crucially, our analysis reveals fundamental bottlenecks inherent in current MLLMs architectures, notably deficient inter-video context retention and poor disambiguation of overlapping entities. CVBench establishes a rigorous framework for advancing pattern recognition methodologies in multi-video scenarios, providing architectural insights for next-generation models. The data and evaluation code are available at: https://github.com/Hokhim2/CVBench.

多视频推理基准测试大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。