让视频推理模型像人一样动态调用工具,逐步分析画面
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
- 模型可随时调用视觉工具,边看边思考
- 在长视频任务上准确率显著提升
- 适合需要多步视觉推理的研究者
视频推理是检验模型能力的综合性测试,要求强大的感知与理解能力。现有方法依赖文本链式思维,常因表示不匹配和感知能力有限而效果不佳。为此,我们提出Weaver——一个端到端可训练的多模态推理智能体系统。Weaver使策略模型能在推理过程中动态调用多种工具,逐步获取关键视觉线索,构建真实的多模态推理路径。同时,采用无轨迹强化学习算法,让系统自由探索工具使用与组合策略。大量实验表明,Weaver在多个复杂视频推理基准上表现优异,尤其在长视频任务中优势明显。
原文摘要 · Abstract (English)
Video reasoning constitutes a comprehensive assessment of a model's capabilities, as it demands robust perceptual and interpretive skills, thereby serving as a means to explore the boundaries of model performance. While recent research has leveraged text-centric Chain-of-Thought reasoning to augment these capabilities, such approaches frequently suffer from representational mismatch and restricted by limited perceptual acuity. To address these limitations, we propose Weaver, a novel, end-to-end trainable multimodal reasoning agentic system. Weaver empowers its policy model to dynamically invoke diverse tools throughout the reasoning process, enabling progressive acquisition of crucial visual cues and construction of authentic multimodal reasoning trajectories. Furthermore, we integrate a reinforcement learning algorithm to allow the system to freely explore strategies for employing and combining these tools with trajectory-free data. Extensive experiments demonstrate that our system, Weaver, enhances performance on several complex video reasoning benchmarks, particularly those involving long videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。