arXiv:2501.04302cs.CVcs.AI2025-01AAAI被引 9

提出分层马尔可夫结构模型,提升自动驾驶视频理解的时序建模能力。

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving

  • 设计双模块架构:上下文马尔可夫与查询马尔可夫,分别捕捉多粒度时空特征
  • 在风险物体检测任务中较当前最优方法提升5.5% mIoU
  • 可即插即用,适合部署于多模态大模型中的自动驾驶视觉理解

随着多模态大语言模型(MLLMs)的普及,自动驾驶面临新机遇与挑战。其中,多模态视频理解对预测动态场景中的未来行为至关重要。然而,自动驾驶视频常包含复杂的时空运动,限制了现有MLLMs的泛化能力。为此,本文提出一种新型分层马尔可夫适应框架(H-MBA),以应对自动驾驶视频中的复杂运动变化。H-MBA包含两个模块:上下文马尔可夫(C-Mamba)和查询马尔可夫(Q-Mamba)。C-Mamba通过多种结构状态空间模型,有效捕捉不同时间分辨率下的多粒度视频上下文;Q-Mamba将当前帧作为可学习查询,有选择性地融合多粒度视频上下文至查询中。该机制可自适应整合多尺度时间分辨率的视频上下文,增强视频理解能力。通过在MLLMs中采用即插即用范式,H-MBA在自动驾驶多模态视频任务中表现优异,例如在风险物体检测任务中,相比之前最先进方法,mIoU提升5.5%。

原文摘要 · Abstract (English)

With the prevalence of Multimodal Large Language Models(MLLMs), autonomous driving has encountered new opportunities and challenges. In particular, multi-modal video understanding is critical to interactively analyze what will happen in the procedure of autonomous driving. However, videos in such a dynamical scene that often contains complex spatial-temporal movements, which restricts the generalization capacity of the existing MLLMs in this field. To bridge the gap, we propose a novel Hierarchical Mamba Adaptation (H-MBA) framework to fit the complicated motion changes in autonomous driving videos. Specifically, our H-MBA consists of two distinct modules, including Context Mamba (C-Mamba) and Query Mamba (Q-Mamba). First, C-Mamba contains various types of structure state space models, which can effectively capture multi-granularity video context for different temporal resolutions. Second, Q-Mamba flexibly transforms the current frame as the learnable query, and attentively selects multi-granularity video context into query. Consequently, it can adaptively integrate all the video contexts of multi-scale temporal resolutions to enhance video understanding. Via a plug-and-play paradigm in MLLMs, our H-MBA shows the remarkable performance on multi-modal video tasks in autonomous driving, e.g., for risk object detection, it outperforms the previous SOTA method with 5.5% mIoU improvement.

视频理解自动驾驶马尔可夫模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。