arXiv:2505.15192cs.CV2025-05被引 3

用动态图结构融合视觉语言模型,提升双手精细动作识别准确率

Leveraging Foundation Models for Multimodal Graph-Based Action Recognition

  • 构建可变的多模态图,节点含帧、物体与文本,边编码时空语义关系
  • 在多个数据集上超越现有方法,关键任务准确率显著提升
  • 适合做视频理解、人机交互等需要细粒度动作分析的研究者

基础模型为多模态视频理解开启了新纪元,能够提取丰富的时空与语义表征。本文提出一种基于图的新框架,结合视觉-语言基础模型:使用 VideoMAE 进行动态视觉编码,BERT 实现上下文文本嵌入,以解决细粒度双侧操作动作识别难题。不同于传统静态图结构,本方法构建自适应多模态图,其中节点代表帧、物体和文本注释,边编码空间、时间与语义关系,图结构随学习到的交互动态演化,实现灵活且上下文感知的推理。图注意力网络中引入任务特定注意力机制,根据动作语义调节边的重要性。在多个基准数据集上的广泛评估表明,该方法持续优于当前最优基线,验证了将基础模型与动态图推理结合在鲁棒且可泛化动作识别中的有效性。

原文摘要 · Abstract (English)

Foundation models have ushered in a new era for multimodal video understanding by enabling the extraction of rich spatiotemporal and semantic representations. In this work, we introduce a novel graph-based framework that integrates a vision-language foundation, leveraging VideoMAE for dynamic visual encoding and BERT for contextual textual embedding, to address the challenge of recognizing fine-grained bimanual manipulation actions. Departing from conventional static graph architectures, our approach constructs an adaptive multimodal graph where nodes represent frames, objects, and textual annotations, and edges encode spatial, temporal, and semantic relationships. These graph structures evolve dynamically based on learned interactions, allowing for flexible and context-aware reasoning. A task-specific attention mechanism within a Graph Attention Network further enhances this reasoning by modulating edge importance based on action semantics. Through extensive evaluations on diverse benchmark datasets, we demonstrate that our method consistently outperforms state-of-the-art baselines, underscoring the strength of combining foundation models with dynamic graph-based reasoning for robust and generalizable action recognition.

动作识别多模态图神经网络视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。