用音视频与场景描述联合建模,提升动作检测精度
JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
- 以角色为中心融合音视频与场景描述信息
- 在三个基准上达到新最好效果,最高提升4.2%
- 适合需要多模态动作理解的视觉研究者
视频动作检测(VAD)旨在定位并分类视频中的动作实例,其本质包含音频、视觉线索和周围场景上下文等多源信息。有效利用这些多模态信息是重大挑战,因模型需精准识别与动作相关的关键线索。本文提出首个融合音频、视觉特征与大容量图像描述模型生成的场景描述上下文的VAD架构——JoVALE。核心在于以角色为中心聚合音视频与场景描述信息,实现对每个角色动作的自适应特征融合。我们设计了基于Transformer的演员中心多模态融合网络,专门捕捉演员间及其多模态上下文的动态交互。在AVA、UCF101-24和JHMDB51-21三个主流基准上的评估表明,引入多模态信息显著提升性能,创下新的最佳纪录。
原文摘要 · Abstract (English)
Video Action Detection (VAD) entails localizing and categorizing action instances within videos, which inherently consist of diverse information sources such as audio, visual cues, and surrounding scene contexts. Leveraging this multi-modal information effectively for VAD poses a significant challenge, as the model must identify action-relevant cues with precision. In this study, we introduce a novel multi-modal VAD architecture, referred to as the Joint Actor-centric Visual, Audio, Language Encoder (JoVALE). JoVALE is the first VAD method to integrate audio and visual features with scene descriptive context sourced from large-capacity image captioning models. At the heart of JoVALE is the actor-centric aggregation of audio, visual, and scene descriptive information, enabling adaptive integration of crucial features for recognizing each actor's actions. We have developed a Transformer-based architecture, the Actor-centric Multi-modal Fusion Network, specifically designed to capture the dynamic interactions among actors and their multi-modal contexts. Our evaluation on three prominent VAD benchmarks, including AVA, UCF101-24, and JHMDB51-21, demonstrates that incorporating multi-modal information significantly enhances performance, setting new state-of-the-art performances in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。