构建视频多模态深度推理新基准,挑战模型跨帧推理能力
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
- 设计长程多帧推理任务,要求模型跨越远距离找证据
- 最先进模型仅达64.3%准确率,推理提升有限
- 适合研究多模态推理、视频理解与模型可解释性的学者
视频的时序结构给多模态大语言模型(MLLMs)定位多帧证据和进行多模态推理带来挑战。现有视频基准多聚焦于理解任务,仅需匹配问题提及的帧(称作'问题帧')并感知少量邻近帧。为填补这一空白,我们提出MMR-V:面向视频中多模态深度推理的基准。该基准具备四大特征:(1) 长程多帧推理:要求模型推断并分析可能远离问题帧的证据帧;(2) 超越感知:问题无法仅通过直接感知回答,需推理隐藏信息;(3) 可靠性:所有任务经人工标注,参考真实用户理解以契合普遍认知;(4) 混淆性:精心设计干扰项标注策略,减少模型捷径依赖。MMR-V包含317个视频和1,257个任务。实验显示当前模型仍难以完成多模态推理,即使最优模型Gemini-2.5-pro也仅达64.3%准确率。此外,当前推理增强策略(如思维链和扩大测试时计算量)带来的提升有限。错误分析表明,多模态推理所需的思维链不同于纯文本推理,部分解释了性能提升受限的原因。我们希望MMR-V能推动多模态推理能力的进一步研究。
原文摘要 · Abstract (English)
The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence and conduct multimodal reasoning. However, existing video benchmarks mainly focus on understanding tasks, which only require models to match frames mentioned in the question (hereafter referred to as "question frame") and perceive a few adjacent frames. To address this gap, we propose MMR-V: A Benchmark for Multimodal Deep Reasoning in Videos. The benchmark is characterized by the following features. (1) Long-range, multi-frame reasoning: Models are required to infer and analyze evidence frames that may be far from the question frame. (2) Beyond perception: Questions cannot be answered through direct perception alone but require reasoning over hidden information. (3) Reliability: All tasks are manually annotated, referencing extensive real-world user understanding to align with common perceptions. (4) Confusability: Carefully designed distractor annotation strategies to reduce model shortcuts. MMR-V consists of 317 videos and 1,257 tasks. Our experiments reveal that current models still struggle with multi-modal reasoning; even the best-performing model, Gemini-2.5-pro, achieves only 64.3% accuracy. Additionally, current reasoning enhancement strategies (Chain-of-Thought and scaling test-time compute) bring limited gains. Error analysis indicates that the CoT demanded for multi-modal reasoning differs from it in textual reasoning, which partly explains the limited performance gains. We hope that MMR-V can inspire further research into enhancing multi-modal reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。