自适应融合多模态数据,提升动作识别准确率与鲁棒性
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
- 采用门控机制动态融合RGB、光流、音频和深度信息
- 在多个基准数据集上显著优于传统单模态方法
- 适合智能安防、人机交互等需要多模态感知的场景
本研究提出一种基于深度神经网络与自适应融合策略的多模态人体动作识别新方法,整合RGB图像、光流、音频及深度信息。通过门控机制实现多模态信息的动态选择性融合,突破传统单模态方法的局限。系统评估多种门控融合策略,确定最优方案,在动作识别、暴力行为检测及自监督学习任务中均取得显著性能提升。该方法有效提取关键特征,构建更完整的动作表征,大幅增强识别准确性与鲁棒性。实验在多个公开基准数据集上验证了其优越性,为监控、人机交互及主动辅助生活等领域提供关键技术支撑。
原文摘要 · Abstract (English)
This study introduces a pioneering methodology for human action recognition by harnessing deep neural network techniques and adaptive fusion strategies across multiple modalities, including RGB, optical flows, audio, and depth information. Employing gating mechanisms for multimodal fusion, we aim to surpass limitations inherent in traditional unimodal recognition methods while exploring novel possibilities for diverse applications. Through an exhaustive investigation of gating mechanisms and adaptive weighting-based fusion architectures, our methodology enables the selective integration of relevant information from various modalities, thereby bolstering both accuracy and robustness in action recognition tasks. We meticulously examine various gated fusion strategies to pinpoint the most effective approach for multimodal action recognition, showcasing its superiority over conventional unimodal methods. Gating mechanisms facilitate the extraction of pivotal features, resulting in a more holistic representation of actions and substantial enhancements in recognition performance. Our evaluations across human action recognition, violence action detection, and multiple self-supervised learning tasks on benchmark datasets demonstrate promising advancements in accuracy. The significance of this research lies in its potential to revolutionize action recognition systems across diverse fields. The fusion of multimodal information promises sophisticated applications in surveillance and human-computer interaction, especially in contexts related to active assisted living.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。