arXiv:2409.11513cs.CVcs.AI2024-09被引 9

用提问引导模型学习动作,提升长时依赖捕捉能力。

Mamba Fusion: Learning Actions Through Questioning

  • 采用共享状态转移矩阵融合视觉与语言模态,降低计算开销。
  • 在Epic-Kitchens-100上达到当前最佳动作识别性能。
  • 通过问答任务聚焦关键动作线索,适合视频理解研究者。

视频语言模型(VLMs)在跨任务泛化和利用语言提示增强学习方面至关重要。尽管基于Transformer的架构已成为视觉-语言训练的主流,但其面临二次计算复杂度、高显存占用以及长时依赖建模困难等问题。为此,我们提出MambaVL,一种利用选择性状态空间模态融合最新进展的新模型,可高效捕捉长程依赖并学习视觉与语言数据的联合表示。MambaVL在双模态间共享状态转移矩阵,使模型能从场景中多视角捕捉动作信息。此外,我们设计了一种问答任务,引导模型关注相关线索。这些问题提供关于动作、物体及环境上下文的关键信息,显著提升性能。结果表明,MambaVL在Epic-Kitchens-100数据集上的动作识别任务中达到当前最优表现,并优于基线方法在动作预测任务中的性能。

原文摘要 · Abstract (English)

Video Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning. While transformer-based architectures have been the de facto in vision-language training, they face challenges like quadratic computational complexity, high GPU memory usage, and difficulty with long-term dependencies. To address these limitations, we introduce MambaVL, a novel model that leverages recent advancements in selective state space modality fusion to efficiently capture long-range dependencies and learn joint representations for vision and language data. MambaVL utilizes a shared state transition matrix across both modalities, allowing the model to capture information about actions from multiple perspectives within the scene. Furthermore, we propose a question-answering task that helps guide the model toward relevant cues. These questions provide critical information about actions, objects, and environmental context, leading to enhanced performance. As a result, MambaVL achieves state-of-the-art performance in action recognition on the Epic-Kitchens-100 dataset and outperforms baseline methods in action anticipation.

视频理解多模态动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。