通过多轮迭代推理,提升长视频理解的准确性和效率。
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
- 采用多轮逐步选段与问题理解机制,动态优化推理过程。
- 在VideoMME、MLVU等数据集上超越现有方法,准确率显著提升。
- 无需外部视觉语言模型,支持端到端训练,适合长视频任务研究者。
长视频理解面临长时依赖和多事件交织的挑战。现有方法多依赖静态推理或外部视觉语言模型(VLMs),存在复杂度高、性能不佳等问题,且缺乏端到端训练。本文提出Video-MTR,一种强化的多轮推理框架,实现关键视频片段的迭代选择与问题理解。不同于单次生成预测的传统流程,Video-MTR通过多轮逐步分析,基于已处理片段的演化理解与当前问题,动态选取视频段落,实现更精细、上下文敏感的分析。为保障中间推理质量,设计新颖的门控双层奖励机制,结合基于答案正确性的轨迹级奖励与强调帧-查询相关性的回合级奖励,联合优化片段选择与问题理解。该机制无需外部VLM,支持端到端训练。在VideoMME、MLVU和EgoSchema等基准上大量实验表明,Video-MTR在准确率与效率上均优于现有方法,推动了长视频理解的最新进展。
原文摘要 · Abstract (English)
Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often rely on static reasoning or external visual-language models (VLMs), which face issues like complexity and sub-optimal performance due to the lack of end-to-end training. In this paper, we propose Video-MTR, a reinforced multi-turn reasoning framework designed to enable iterative key video segment selection and question comprehension. Unlike traditional video reasoning pipeline, which generate predictions in a single turn, Video-MTR performs reasoning in multiple turns, selecting video segments progressively based on the evolving understanding of previously processed segments and the current question. This iterative process allows for a more refined and contextually aware analysis of the video. To ensure intermediate reasoning process, we introduce a novel gated bi-level reward system, combining trajectory-level rewards based on answer correctness and turn-level rewards emphasizing frame-query relevance. This system optimizes both video segment selection and question comprehension, eliminating the need for external VLMs and allowing end-to-end training. Extensive experiments on benchmarks like VideoMME, MLVU, and EgoSchema demonstrate that Video-MTR outperforms existing methods in both accuracy and efficiency, advancing the state-of-the-art in long video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。