让多个智能体协作理解长视频,不丢信息还省算力。
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

- 分段处理视频,每个智能体专注局部,共享紧凑特征
- 在相同算力下,比现有模型更准确理解长视频
- 适合需要高效处理长视频的场景,如监控、教育
多模态大语言模型在长视频任务中受限于固定的感知上下文预算。现有代理方法依赖规则预处理,常导致信息丢失、成本高且依赖文本中间表示。我们提出MACF——一种端到端的多智能体协作框架,将各智能体的感知预算与全局视频复杂度解耦,实现可扩展的视频理解并保持视觉保真度。MACF将视频分段,由局部预算智能体分别处理,并通过原生潜在通信协议实现整体推理。每个智能体将局部观察编码为共享嵌入空间中的紧凑、任务充分的标记,由中心协调器高效协作。我们引入课程训练策略,逐步强化语义对齐、证据摘要与跨智能体协同。在多种视频理解基准上的实验表明,MACF在相同预算约束下持续优于最先进多模态大模型和多代理系统,验证了潜在协作在可扩展视频理解中的有效性。
原文摘要 · Abstract (English)
Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existing agentic methods mitigate this via rule-based preprocessing, yet often suffer from information loss, high cost, and reliance on textual intermediates. We propose MACF, an end-to-end Multi-Agent Collaboration Framework that decouples per-agent perception budgets from global video complexity, enabling scalable video understanding while preserving visual fidelity. MACF partitions videos into segments for locally budgeted agents and enables holistic reasoning via an agent-native latent communication protocol. Each agent encodes partial observations into compact, task-sufficient tokens in a shared embedding space, allowing efficient and information-preserving collaboration by a central coordinator. We introduce a curriculum training strategy that progressively enforces semantic alignment, evidence summarization, and cross-agent coordination. Extensive experiments on diverse video understanding benchmarks show that MACF consistently outperforms state-of-the-art MLLMs and multi-agent systems under identical budget constraints, demonstrating the effectiveness of our latent collaboration for scalable video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。