通过分步推理实现视频从感知到认知的深度理解
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

- 构建像素级时空感知模型MotionEpic,融合场景图表示
- 提出VoT框架,分步完成从像素到认知的视频推理
- 首次实现类人视频推理,在多个基准上显著超越现有方法
现有视频理解研究在复杂视频中仍难以实现深层认知与推理,主要受限于两个关键瓶颈:细粒度时空感知理解和认知层面的场景理解。本文提出新解决方案。首先引入新型多模态大模型MotionEpic,通过整合视频时空场景图(STSG)表示,实现像素级时空定位。在此基础上,构建视频思维链(VoT)推理框架,继承思维链(CoT)核心思想,将复杂任务拆解为可管理的子问题,从低层像素感知逐步推进至高层认知解释。在多个复杂视频问答基准上的大量实验表明,该框架显著提升现有最先进水平。据我们所知,这是首次成功将思维链技术应用于实现人类级视频推理,展现出向更广泛视频理解场景扩展的巨大潜力。项目开源地址:https://haofei.vip/VoT
原文摘要 · Abstract (English)
Existing research of video understanding still struggles to achieve in-depth comprehension and reasoning in complex videos, primarily due to the under-exploration of two key bottlenecks: fine-grained spatial-temporal perceptive understanding and cognitive-level video scene comprehension. This paper bridges the gap by presenting a novel solution. We first introduce a novel video Multimodal Large Language Model (MLLM), MotionEpic, which achieves fine-grained pixel-level spatial-temporal video grounding by integrating video spatial-temporal scene graph (STSG) representation. Building upon MotionEpic, we then develop a Video-of-Thought (VoT) reasoning framework. VoT inherits the Chain-of-Thought (CoT) core, breaking down a complex task into simpler and manageable sub-problems, and addressing them step-by-step from a low-level pixel perception to high-level cognitive interpretation. Extensive experiments across various complex video QA benchmarks demonstrate that our overall framework strikingly boosts existing state-of-the-art. To our knowledge, this is the first attempt at successfully implementing the CoT technique for achieving human-level video reasoning, where we show great potential in extending it to a wider range of video understanding scenarios. Project is open at https://haofei.vip/VoT
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。