arXiv:2606.27828cs.CV2026-06被引 3

构建可控视频逻辑推理测试集,评估模型跨帧推理能力。

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

论文配图:Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
图 1 · 摘自论文原文
  • 设计五类受控逻辑操作,分离时序推理与场景复杂度
  • 25个细粒度任务,支持难度可调与中间状态诊断
  • 揭示大模型在复杂时序推理上的显著差距,适合研究多模态推理

多模态大模型能否基于动态视觉证据进行推理,而非仅识别单帧中的物体或事件?我们称此能力为视频时序逻辑推理,要求模型在帧间演变中持续、更新并组合视觉证据。现有视频基准常将此能力与场景复杂度、静态识别或非控制的时间变化混杂。为此,我们提出Video-MME-Logical,一个围绕五种时序逻辑操作(状态追踪、顺序计数、时间排序、动态空间性、结构组合)构建的受控基准。该基准包含25个精细任务类别,通过受控的物体状态、转换、时间依赖和逻辑组合生成。支持通过调整时间跨度和推理复杂度实现难度可控的最终答案评估,并可通过验证模型是否恢复所需逻辑推理轨迹来实现中间状态诊断。对先进多模态大模型的实验显示,人类与模型之间存在显著差距,尤其在时序逻辑复杂度增加时。在最多50万条生成样本上进行监督微调虽能提升性能,但仍不足以弥合推理差距,使Video-MME-Logical成为分析和改进多模态大模型时序逻辑推理能力的可扩展测试平台。

原文摘要 · Abstract (English)

Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events in individual frames? This ability, which we refer to as video temporal-logical reasoning, requires models to maintain, update, and compose evidence as visual states evolve across frames. Existing video benchmarks often conflate this capability with scene complexity, static recognition, or uncontrolled temporal variation. To isolate this capability, we introduce Video-MME-Logical, a controlled benchmark organized around five temporal-logical operations: state tracking, sequential counting, temporal ordering, dynamic spatiality, and structural composition. The benchmark contains 25 fine-grained task categories generated with controlled object states, transitions, temporal dependencies, and logical compositions. It enables difficulty-controlled final-answer evaluation by varying temporal horizon and reasoning complexity, and supports intermediate-state diagnostics by verifying whether models recover the required logical reasoning trace before producing the final answer. Experiments with state-of-the-art MLLMs reveal a substantial human-model gap, especially as temporal-logical complexity increases. Supervised fine-tuning on up to 500K generated samples improves performance but remains insufficient to close the reasoning gap, positioning Video-MME-Logical as a scalable testbed for analyzing and improving temporal-logical reasoning in MLLMs.

多模态推理视频理解逻辑推理评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。