arXiv:2504.17213cs.CVcs.AI2025-04被引 6

通过分层注意力聚焦与自我反思,提升智能体视频理解的准确率。

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding

  • 采用多模态粗细粒度相关性感知,精准定位视频中与问题相关的片段。
  • 在EgoSchema上比领先方法提升5%,在多个数据集上超越当前最佳表现。
  • 适合需要高精度长时视频理解的智能体应用,如自动驾驶、机器人交互。

尽管大模型发展迅速,视频理解仍是极具挑战的任务。相比文本或图像,视频常包含冗余信息,需大模型在全局层面合理分配注意力以实现全面准确的理解。为此,我们提出一种面向智能体视频理解的多模态层次注意力聚焦自我反思推理框架(MASR)。其核心创新在于能检测并优先处理与查询高度相关的视频片段。首先,MASR通过多模态粗细粒度相关性感知(MCRS)增强获取的上下文信息与查询之间的关联性;其次,采用膨胀时间扩展(DTE)机制,在基于MCRS聚焦的帧中提取语义信息时降低遗漏关键细节的风险。通过在自我反思推理过程中迭代应用MCRS与DTE,MASR可自适应调整注意力,提取高度相关上下文,从而提升回答准确性。在EgoSchema数据集上,MASR相较先前领先方法提升5%;在Next-QA和IntentQA数据集上分别优于现有最佳水平0.2%和0.3%;在包含长时视频的Video-MME数据集上,也优于其他基于智能体的方法。

原文摘要 · Abstract (English)

Even in the era of rapid advances in large models, video understanding remains a highly challenging task. Compared to texts or images, videos commonly contain more information with redundancy, requiring large models to properly allocate attention at a global level for comprehensive and accurate understanding. To address this, we propose a Multimodal hierarchical Attention focusing Self-reflective Reasoning (MASR) framework for agent-based video understanding. The key innovation lies in its ability to detect and prioritize segments of videos that are highly relevant to the query. Firstly, MASR realizes Multimodal Coarse-to-fine Relevance Sensing (MCRS) which enhances the correlation between the acquired contextual information and the query. Secondly, MASR employs Dilated Temporal Expansion (DTE) to mitigate the risk of missing crucial details when extracting semantic information from the focused frames selected through MCRS. By iteratively applying MCRS and DTE in the self-reflective reasoning process, MASR is able to adaptively adjust the attention to extract highly query-relevant context and therefore improve the response accuracy. In the EgoSchema dataset, MASR achieves a remarkable 5% performance gain over previous leading approaches. In the Next-QA and IntentQA datasets, it outperforms the state-of-the-art standards by 0.2% and 0.3% respectively. In the Video-MME dataset that contains long-term videos, MASR also performs better than other agent-based methods.

视频理解多模态注意力机制智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。