arXiv:2511.19524cs.CVcs.MA2025-11中稿 · CVPR被引 7

多智能体协作规划让视频理解更准,动态调整策略应对复杂视频。

VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning

  • 多个智能体各自生成工具调用策略,协同探索视频内容。
  • 在八个基准上达顶尖水平,长视频任务上超越GPT-4o 15.6%。
  • 适合需要深度时序分析的视频理解场景,如监控、教育视频分析。

通过引入工具增强的多模态大语言模型,多智能体框架正在推动视频理解的发展。然而,现有方法大多采用静态且不可学习的工具调用机制,限制了对时空复杂视频中多样线索的发现。为此,我们提出一种新型多智能体系统VideoChat-M1,采用独特的协同策略规划(CPP)范式,包含三个关键过程:(1) 策略生成:每个智能体根据用户查询生成专属的工具调用策略;(2) 策略执行:各智能体依次调用相关工具,探索视频内容;(3) 策略通信:在执行过程中,智能体相互交流,动态更新自身策略。通过该协作框架,所有智能体协同工作,基于同伴提供的上下文信息持续优化策略,以有效响应用户查询。此外,我们结合简洁的多智能体强化学习(MARL)方法,使策略智能体团队可联合优化,同时受到最终答案奖励和中间协作反馈的引导。大量实验表明,VideoChat-M1在涵盖四个任务的八个基准上达到最先进性能。尤其在LongVideoBench上,优于SOTA模型Gemini 2.5 Pro 3.6%,优于GPT-4o 15.6%。

原文摘要 · Abstract (English)

By leveraging tool-augmented Multimodal Large Language Models (MLLMs), multi-agent frameworks are driving progress in video understanding. However, most of them adopt static and non-learnable tool invocation mechanisms, which limit the discovery of diverse clues essential for robust perception and reasoning regarding temporally or spatially complex videos. To address this challenge, we propose a novel Multi-agent system for video understanding, namely VideoChat-M1. Instead of using a single or fixed policy, VideoChat-M1 adopts a distinct Collaborative Policy Planning (CPP) paradigm with multiple policy agents, which comprises three key processes. (1) Policy Generation: Each agent generates its unique tool invocation policy tailored to the user's query; (2) Policy Execution: Each agent sequentially invokes relevant tools to execute its policy and explore the video content; (3) Policy Communication: During the intermediate stages of policy execution, agents interact with one another to update their respective policies. Through this collaborative framework, all agents work in tandem, dynamically refining their preferred policies based on contextual insights from peers to effectively respond to the user's query. Moreover, we equip our CPP paradigm with a concise Multi-Agent Reinforcement Learning (MARL) method. Consequently, the team of policy agents can be jointly optimized to enhance VideoChat-M1's performance, guided by both the final answer reward and intermediate collaborative process feedback. Extensive experiments demonstrate that VideoChat-M1 achieves SOTA performance across eight benchmarks spanning four tasks. Notably, on LongVideoBench, our method outperforms the SOTA model Gemini 2.5 pro by 3.6% and GPT-4o by 15.6%.

视频理解多智能体强化学习长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。