让大模型像看图思考一样自主分析视频,无需外部工具。
Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
- 用强化学习激发模型自动生成视频线索进行推理
- 在多个视频推理数据集上超越现有7B模型表现
- 适合想提升视频理解能力的研究者与开发者
近期图像推理方法,特别是“看图思考”范式,在多模态大模型中取得显著成功;然而这一动态推理机制尚未拓展至视频任务。本文提出 Video-Thinker,通过激活多模态大模型内在的“定位”与“描述”能力,在推理过程中自主生成推理线索,实现视频思考。为此,我们构建了 Video-Thinker-10K 数据集,包含链式思维中自主使用工具的序列。训练策略采用监督微调(SFT)学习推理格式,再通过组相对策略优化(GRPO)强化推理能力。该方法使模型能自主完成视频定位与描述任务,无需依赖外部工具。大量实验表明,Video-Thinker 在域内任务及挑战性跨域视频推理基准(如 Video-Holmes、CG-Bench-Reasoning、VRBench)上均取得显著提升,其 7B 版本大幅超越 Video-R1 等基线,成为 7B 级别模型中的最先进水平。
原文摘要 · Abstract (English)
Recent advances in image reasoning methods, particularly "Thinking with Images", have demonstrated remarkable success in Multimodal Large Language Models (MLLMs); however, this dynamic reasoning paradigm has not yet been extended to video reasoning tasks. In this paper, we propose Video-Thinker, which empowers MLLMs to think with videos by autonomously leveraging their intrinsic "grounding" and "captioning" capabilities to generate reasoning clues throughout the inference process. To spark this capability, we construct Video-Thinker-10K, a curated dataset featuring autonomous tool usage within chain-of-thought reasoning sequences. Our training strategy begins with Supervised Fine-Tuning (SFT) to learn the reasoning format, followed by Group Relative Policy Optimization (GRPO) to strengthen this reasoning capability. Through this approach, Video-Thinker enables MLLMs to autonomously navigate grounding and captioning tasks for video reasoning, eliminating the need for constructing and calling external tools. Extensive experiments demonstrate that Video-Thinker achieves significant performance gains on both in-domain tasks and challenging out-of-domain video reasoning benchmarks, including Video-Holmes, CG-Bench-Reasoning, and VRBench. Our Video-Thinker-7B substantially outperforms existing baselines such as Video-R1 and establishes state-of-the-art performance among 7B-sized MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。