用动态特征选择提升实时人体动作识别准确率与效率
A Novel Deep Hybrid Framework with Ensemble-Based Feature Optimization for Robust Real-Time Human Activity Recognition
- 融合Inception-V3与注意力LSTM捕捉空间与时序特征
- 仅用7个特征达99.65%准确率,推理更快
- 适合资源受限的实时监控系统应用
实时人体动作识别(HAR)在智能环境、公共安全、辅助技术及自主监控中应用广泛。现有系统存在可扩展性差、计算开销高等问题,主要源于冗余特征。本文定制Inception-V3模型,结合区域与边界感知操作,分别使用平均池化和最大池化增强区域一致性、抑制噪声并捕获判别性局部特征,通过下采样提升鲁棒性。为有效编码运动动态,采用注意力增强型长短期记忆网络(AA-LSTM)学习帧间时序依赖。视频数据提取特征后,通过一种新型动态复合特征选择方法——自适应动态适应共享与注意力(ADFSA)进行优化。该机制嵌入遗传算法,动态平衡精度、冗余降低、特征唯一性与复杂度最小化,选出紧凑且具判别力的特征子集。结果表明,仅需7个特征即可在具有遮挡、背景杂乱、复杂运动和光照不良的UCF-YouTube数据集上实现最高99.65%准确率,并显著提升推理速度。
原文摘要 · Abstract (English)
Real-time Human Activity Recognition (HAR) has wide-ranging applications in areas such as context-aware environments, public safety, assistive technologies, and autonomous monitoring and surveillance systems. However, existing real-time HAR systems face significant challenges, including limited scalability and high computational costs arising from redundant features. To address these issues, the Inception-V3 model was customized with region-based and boundary-aware operations, using average pooling and max pooling, respectively, to enhance region homogeneity, suppress noise, and capture discriminative local features, while improving robustness through down-sampling. Furthermore, to effectively encode motion dynamics, an Attention-Augmented Long Short-Term Memory (AA-LSTM) network was employed to learn temporal dependencies across video frames. Features are extracted from video dataset and are then optimized through a novel proposed dynamic composite feature selection method called Adaptive Dynamic Fitness Sharing and Attention (ADFSA). This ADFSA mechanism is embedded within a genetic algorithm to select a compact, optimized subset of features by dynamically balancing multiple objectives, accuracy, redundancy reduction, feature uniqueness, and complexity minimization. As a result, the selected subset of diverse and discriminative features enables lightweight machine learning classifiers to achieve accurate and robust HAR in heterogeneous environments. Experimental results demonstrate up to 99.65\% accuracy using as few as seven selected features, with improved inference time on the challenging UCF-YouTube dataset, which includes factors such as occlusion, cluttered backgrounds, complex motion dynamics, and poor illumination conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。