双阶段特征融合提升人体动作识别准确率
Multilevel neural networks with dual-stage feature fusion for human activity recognition

- 分两层设计网络,结合中间与后期特征融合
- 在两个公开数据集上准确率超越单一融合方法
- 适合需要高精度动作识别的应用场景
人体动作识别(HAR)利用传感器数据识别人类行为。卷积神经网络(CNN)、长短期记忆网络(LSTM)、卷积LSTM及其混合模型在多个领域表现优异。本研究提出一种两级网络架构,包含双阶段特征融合:晚期融合(合并第一层输出)与中期融合(整合第一、二层特征)。我们评估了15种不同组合的CNN、LSTM和卷积LSTM架构,对比含与不含中期融合的晚期融合方案,以确定最优配置。在两个公开基准数据集上的实验表明,同时采用晚期与中期融合的架构,性能优于仅用晚期融合的方法。最优配置显著优于基线模型,验证了其在HAR中的有效性。
原文摘要 · Abstract (English)
Human activity recognition (HAR) refers to the process of identifying human actions and activities using data collected from sensors. Neural networks, such as convolutional neural networks (CNNs), long short-term memory (LSTM) networks, convolutional LSTM, and their hybrid combinations, have demonstrated exceptional performance in various research domains. Developing a multilevel individual or hybrid model for HAR involves strategically integrating multiple networks to capitalize on their complementary strengths. The structural arrangement of these components is a critical factor influencing the overall performance. This study explores a novel framework of a two-level network architecture with dual-stage feature fusion: late fusion, which combines the outputs from the first network level, and intermediate fusion, which integrates the features from both the first and second levels. We evaluated $15$ different network architectures of CNNs, LSTMs, and convolutional LSTMs, incorporating late fusion with and without intermediate fusion, to identify the optimal configuration. Experimental evaluation on two public benchmark datasets demonstrates that architectures incorporating both late and intermediate fusion achieve higher accuracy than those relying on late fusion alone. Moreover, the optimal configuration outperforms baseline models, thereby validating its effectiveness for HAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。