arXiv:2511.12263cs.CVcs.AI2025-11AAAI被引 3

首个评估多视频推理能力的基准,揭示大模型在跨视频理解上的短板。

CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models

  • 构建多层级任务体系,覆盖真实场景下的跨视频推理需求。
  • 包含5331段视频与9015个问答对,测试多种模型跨视频分析能力。
  • 发现当前大模型普遍难以整合多视频信息,推动未来改进方向。

跨视频推理(CVR)是视频理解中的重大挑战,要求模型同时理解多段视频以聚合和比较信息。现有视频理解基准多聚焦单视频分析,无法有效评估多模态大语言模型(MLLMs)在多视频场景下的推理能力。尽管近期研究提出针对多视角视频的评测,但任务类型有限,难以全面反映真实世界的复杂性。为此,我们提出CrossVid,首个系统评估MLLM在跨视频情境下时空推理能力的基准。CrossVid涵盖四个高层维度与十项具体任务,高度模拟现实场景;提供5,331段视频及9,015个具有挑战性的问答对,支持单选、多选与开放问答。在多个开源与闭源模型上进行实验发现,Gemini-2.5-Pro表现最佳,平均准确率达50.4%。深入案例研究表明,当前多数模型在跨视频推理中表现不佳,主要因无法有效整合或对比分布在多视频中的证据。这些发现凸显CrossVid在推动未来提升MLLM跨视频推理能力方面的潜力。

原文摘要 · Abstract (English)

Cross-Video Reasoning (CVR) presents a significant challenge in video understanding, which requires simultaneous understanding of multiple videos to aggregate and compare information across groups of videos. Most existing video understanding benchmarks focus on single-video analysis, failing to assess the ability of multimodal large language models (MLLMs) to simultaneously reason over various videos. Recent benchmarks evaluate MLLMs' capabilities on multi-view videos that capture different perspectives of the same scene. However, their limited tasks hinder a thorough assessment of MLLMs in diverse real-world CVR scenarios. To this end, we introduce CrossVid, the first benchmark designed to comprehensively evaluate MLLMs' spatial-temporal reasoning ability in cross-video contexts. Firstly, CrossVid encompasses a wide spectrum of hierarchical tasks, comprising four high-level dimensions and ten specific tasks, thereby closely reflecting the complex and varied nature of real-world video understanding. Secondly, CrossVid provides 5,331 videos, along with 9,015 challenging question-answering pairs, spanning single-choice, multiple-choice, and open-ended question formats. Through extensive experiments on various open-source and closed-source MLLMs, we observe that Gemini-2.5-Pro performs best on CrossVid, achieving an average accuracy of 50.4%. Notably, our in-depth case study demonstrates that most current MLLMs struggle with CVR tasks, primarily due to their inability to integrate or compare evidence distributed across multiple videos for reasoning. These insights highlight the potential of CrossVid to guide future advancements in enhancing MLLMs' CVR capabilities.

视频理解多模态推理能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。