arXiv:2505.23359cs.CV2025-05被引 20

新基准测试揭示多模态模型在复杂视觉推理中的短板。

VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?

  • 设计视觉主导的视频推理评测集,要求逐步推理解析隐含状态。
  • 多数模型表现不佳,最强模型仅56%准确率,需长链思维才有效。
  • 适合研究视频理解、长链推理与多模态模型能力的学者参考。

近期研究表明,长链思维(CoT)可显著提升大语言模型在复杂任务上的表现。然而,这一优势尚未在视频理解领域得到验证,因现有基准缺乏足够推理深度。尽管已有研究提出视频推理基准,但任务多依赖知识而非视觉内容。为此,我们提出VideoReasonBench,一个专为评估视觉主导的复杂视频推理而设计的基准。每个视频描述一个潜在状态上的细粒度操作序列,仅部分可见。问题涵盖三个递进层次:回忆观察到的视觉信息、推断隐含状态内容、预测视频外信息。模型需精确回忆多个操作并逐步推理才能答对。我们全面评估了18个先进多模态大模型(MLLMs),发现大多数在复杂视频推理中表现差,例如GPT-4o仅达6.9%准确率;而思维增强版Gemini-2.5-Pro以56.0%准确率显著领先。进一步分析显示,扩展思维预算在现有视频基准上几乎无效,但在VideoReasonBench上至关重要。

原文摘要 · Abstract (English)

Recent studies have shown that long chain-of-thought (CoT) reasoning can significantly enhance the performance of large language models (LLMs) on complex tasks. However, this benefit is yet to be demonstrated in the domain of video understanding, since most existing benchmarks lack the reasoning depth required to demonstrate the advantages of extended CoT chains. While recent efforts have proposed benchmarks aimed at video reasoning, the tasks are often knowledge-driven and do not rely heavily on visual content. To bridge this gap, we introduce VideoReasonBench, a benchmark designed to evaluate vision-centric, complex video reasoning. To ensure visual richness and high reasoning complexity, each video in VideoReasonBench depicts a sequence of fine-grained operations on a latent state that is only visible in part of the video. The questions evaluate three escalating levels of video reasoning skills: recalling observed visual information, inferring the content of latent states, and predicting information beyond the video. Under such task setting, models have to precisely recall multiple operations in the video, and perform step-by-step reasoning to get correct final answers for these questions. Using VideoReasonBench, we comprehensively evaluate 18 state-of-the-art multimodal LLMs (MLLMs), finding that most perform poorly on complex video reasoning -- e.g., GPT-4o achieves only 6.9% accuracy -- while the thinking-enhanced Gemini-2.5-Pro significantly outperforms others with 56.0% accuracy. Our investigations on "test-time scaling" further reveal that extended thinking budget, while offering none or minimal benefits on existing video benchmarks, is essential for improving the performance on VideoReasonBench.

视频推理多模态模型长链思维基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。