用智能体生成视频推理数据,让大模型更懂复杂视频逻辑。
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
- 设计多智能体框架模拟反复观看,生成带视频依据的推理链条。
- 在5个基准上超越现有方法,平均性能达新高。
- 适合研究视频理解、推理链生成与强化学习对齐的学者。
尽管基于可验证奖励的强化学习(RLVR)显著提升了大型视觉语言模型(LVLMs)的图像推理能力,但其在复杂视频推理中的应用仍不充分。主要瓶颈在于数据匮乏:现有数据集缺乏具有挑战性的多跳问题和高质量、视频锚定的思维链(CoT)数据。为此,我们提出ReWatch,一个大规模视频推理数据集,包含三个组件:ReWatch-Caption、ReWatch-QA 和 ReWatch-CoT。核心创新是多智能体ReAct框架,通过显式建模信息检索与验证,模拟人类“重看”过程生成视频引导的推理轨迹。基于该数据集,我们通过监督微调(SFT)与新的RLVR框架训练出ReWatch-R1。该框架引入观察与推理(O&R)奖励机制,同时评估答案正确性与推理内容与视频的一致性,直接惩罚幻觉。实验表明,ReWatch-R1在五个挑战性视频推理基准上达到当前最优平均表现。
原文摘要 · Abstract (English)
While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasoning remains underdeveloped. This gap stems primarily from a critical data bottleneck: existing datasets lack the challenging, multi-hop questions and high-quality, video-grounded Chain-of-Thought (CoT) data necessary to effectively bootstrap RLVR. To address this, we introduce ReWatch, a large-scale dataset built to foster advanced video reasoning. We propose a novel multi-stage synthesis pipeline to synthesize its three components: ReWatch-Caption, ReWatch-QA, and ReWatch-CoT. A core innovation is our Multi-Agent ReAct framework for CoT synthesis, which simulates a human-like "re-watching" process to generate video-grounded reasoning traces by explicitly modeling information retrieval and verification. Building on this dataset, we develop ReWatch-R1 by post-training a strong baseline LVLM with Supervised Fine-Tuning (SFT) and our RLVR framework. This framework incorporates a novel Observation \& Reasoning (O\&R) reward mechanism that evaluates both the final answer's correctness and the reasoning's alignment with video content, directly penalizing hallucination. Our experiments show that ReWatch-R1 achieves state-of-the-art average performance on five challenging video reasoning benchmarks. Project Page: https://rewatch-r1.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。