arXiv:2604.21444cs.AI2026-04被引 1

提出分层多智能体框架,让视频理解更懂因果与时间逻辑

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration

论文配图:HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration
图 1 · 摘自论文原文
  • 用混合树结构保留视频时间拓扑,按语义聚类关键片段
  • 根据问题生成精准视觉提示,提升对复杂问题的响应能力
  • 动态分配智能体角色,适合需要长程推理的任务场景

长视频理解面临时空冗余与跨时序叙事依赖的挑战。现有结构化表示虽压缩视觉信息,但常牺牲时间连贯性,影响因果推理。当前多智能体框架采用固定流程,难以适配不同问题需求。本文提出HiCrew,一种分层多智能体框架,包含三项核心贡献:第一,提出混合树结构,利用镜头边界检测保留时间拓扑,并在语义一致片段内进行相关性引导的分层聚类;第二,设计问题感知的摘要机制,生成意图驱动的视觉提示,产出面向精度的语义描述;第三,引入规划层,根据问题复杂度动态选择智能体角色与执行路径。在EgoSchema和NExT-QA数据集上的实验表明,该方法在多种问题类型上表现优异,尤其在时间与因果推理任务中因保持层级结构而取得显著提升。

原文摘要 · Abstract (English)

Long-form video understanding remains fundamentally challenged by pervasive spatiotemporal redundancy and intricate narrative dependencies that span extended temporal horizons. While recent structured representations compress visual information effectively, they frequently sacrifice temporal coherence, which is critical for causal reasoning. Meanwhile, existing multi-agent frameworks operate through rigid, pre-defined workflows that fail to adapt their reasoning strategies to question-specific demands. In this paper, we introduce HiCrew, a hierarchical multi-agent framework that addresses these limitations through three core contributions. First, we propose a Hybrid Tree structure that leverages shot boundary detection to preserve temporal topology while performing relevance-guided hierarchical clustering within semantically coherent segments. Second, we develop a Question-Aware Captioning mechanism that synthesizes intent-driven visual prompts to generate precision-oriented semantic descriptions. Third, we integrate a Planning Layer that dynamically orchestrates agent collaboration by adaptively selecting roles and execution paths based on question complexity. Extensive experiments on EgoSchema and NExT-QA validate the effectiveness of our approach, demonstrating strong performance across diverse question types with particularly pronounced gains in temporal and causal reasoning tasks that benefit from our hierarchical structure-preserving design.

视频理解多智能体因果推理长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。