用统一架构融合视觉与骨骼数据,提升动作识别鲁棒性
Unified Framework with Consistency across Modalities for Human Activity Recognition
- 设计可处理时空交互的通用查询机制COMPUTER
- 通过模态一致性损失提升多模态联合建模效果
- 适用于复杂场景下群体动作识别与定位任务
视频中人体活动识别因时空复杂性和上下文依赖性而困难。以往研究多依赖单一模态(如RGB或骨骼数据),难以发挥多模态互补优势。现有方法虽尝试简单特征融合,但受限于模态间表示差异,难以构建统一神经网络有效利用互补信息。为此,本文提出一种全面的多模态框架,核心是新型组合查询机COMPUTER(CompositionaL Human-centric QUery machine),该通用架构可建模目标人与其环境在时空上的交互。得益于其灵活设计,COMPUTER能为不同输入模态提取独特表征。同时引入一致性损失,强制不同模态预测结果一致,从而挖掘多模态输入的互补信息以实现鲁棒的人体运动识别。在动作定位与群体活动识别任务上,实验表明本方法优于当前最优模型。
原文摘要 · Abstract (English)
Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their ability to exploit the complementary advantages across modalities. Recent studies focus on combining these two modalities using simple feature fusion techniques. However, due to the inherent disparities in representation between these input modalities, designing a unified neural network architecture to effectively leverage their complementary information remains a significant challenge. To address this, we propose a comprehensive multimodal framework for robust video-based human activity recognition. Our key contribution is the introduction of a novel compositional query machine, called COMPUTER ($\textbf{COMP}ositional h\textbf{U}man-cen\textbf{T}ric qu\textbf{ER}y$ machine), a generic neural architecture that models the interactions between a human of interest and its surroundings in both space and time. Thanks to its versatile design, COMPUTER can be leveraged to distill distinctive representations for various input modalities. Additionally, we introduce a consistency loss that enforces agreement in prediction between modalities, exploiting the complementary information from multimodal inputs for robust human movement recognition. Through extensive experiments on action localization and group activity recognition tasks, our approach demonstrates superior performance when compared with state-of-the-art methods. Our code is available at: https://github.com/tranxuantuyen/COMPUTER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。