构建视频理解新基准,用反事实推理检验模型逻辑能力。
Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation
- 设计多维反事实视频评测框架,分解复杂问题为子问题
- 发现子问题准确率与反事实推理能力强相关
- 适合评估大模型在动态场景下的推理鲁棒性
反事实推理对鲁棒视频理解至关重要,但现有多模态基准对此关注不足。本文提出 extbf{COVER}( extbf{—}Counterfactual Video Reasoning),一个跨抽象-具体、感知-认知维度的多维多模态评测基准。相比以往基准,COVER 将复杂问题分解为结构化子问题,实现细粒度推理分析。对商用和开源模型的实验表明,子问题准确率与反事实推理性能高度相关,凸显结构化推理在视频理解中的关键作用。结果进一步揭示:提升模型推理能力是增强视频理解鲁棒性的核心。COVER 为评估多模态大模型在动态环境中的逻辑推理能力树立了新标准。代码已公开于 https://github.com/gongyifan-hash/COVER-Benchmark。
原文摘要 · Abstract (English)
Counterfactual reasoning is crucial for robust video understanding but remains underexplored in existing multimodal benchmarks. In this paper, we introduce \textbf{COVER} (\textbf{\underline{CO}}unterfactual \textbf{\underline{V}}id\textbf{\underline{E}}o \textbf{\underline{R}}easoning), a multidimensional multimodal benchmark that systematically evaluates MLLMs across the abstract-concrete and perception-cognition dimensions. Beyond prior multimodal benchmarks, COVER decomposes complex queries into structured sub-questions, enabling fine-grained reasoning analysis. Experiments on commercial and open-source models reveal a strong correlation between sub-question accuracy and counterfactual reasoning performance, highlighting the role of structured inference in video understanding. Furthermore, our results suggest a key insight: enhancing the reasoning capability of models is essential for improving the robustness of video understanding. COVER establishes a new standard for assessing MLLMs' logical reasoning abilities in dynamic environments. Our work is available at https://github.com/gongyifan-hash/COVER-Benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。