arXiv:2603.17307cs.CVcs.AI2026-03中稿 · CVPR被引 4

模仿人类认知的多智能体系统,提升长视频理解能力。

Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding

  • 模拟人脑认知,分解任务并加强协作推理。
  • 在多个基准上达到领先效果,LVBench提升5.0%。
  • 适合需要长时序分析与复杂意图识别的研究者。

尽管多模态大模型代理(MLLM agents)发展迅速且应用广泛,但在高信息密度、长时序跨度的长视频理解(LVU)任务中仍表现不佳。现有研究发现,简单的任务分解与协作机制难以应对长链推理任务;而通过嵌入检索压缩时间上下文,可能丢失复杂问题的关键信息。为此,本文提出Symphony——一种受人类认知启发的多智能体系统,通过将LVU任务细粒度拆解,并引入基于反思的深度推理协作机制,显著增强推理能力。同时,Symphony采用基于视觉语言模型(VLM)的定位方法,分析任务并评估视频片段的相关性,有效提升对隐含意图和长跨度问题的定位能力。实验表明,Symphony在LVBench、LongVideoBench、VideoMME和MLVU等多个基准上均达当前最优性能,其中在LVBench上较之前最优方法提升5.0%。代码已开源:https://github.com/Haiyang0226/Symphony。

原文摘要 · Abstract (English)

Despite rapid developments and widespread applications of MLLM agents, they still struggle with long-form video understanding (LVU) tasks, which are characterized by high information density and extended temporal spans. Recent research on LVU agents demonstrates that simple task decomposition and collaboration mechanisms are insufficient for long-chain reasoning tasks. Moreover, directly reducing the time context through embedding-based retrieval may lose key information of complex problems. In this paper, we propose Symphony, a multi-agent system, to alleviate these limitations. By emulating human cognition patterns, Symphony decomposes LVU into fine-grained subtasks and incorporates a deep reasoning collaboration mechanism enhanced by reflection, effectively improving the reasoning capability. Additionally, Symphony provides a VLM-based grounding approach to analyze LVU tasks and assess the relevance of video segments, which significantly enhances the ability to locate complex problems with implicit intentions and large temporal spans. Experimental results show that Symphony achieves state-of-the-art performance on LVBench, LongVideoBench, VideoMME, and MLVU, with a 5.0% improvement over the prior state-of-the-art method on LVBench. Code is available at https://github.com/Haiyang0226/Symphony.

多智能体长视频理解认知启发视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。