轻量级视频动作识别模型,适配边缘设备且可量化视觉退化影响。
Resource-Efficient RGB-Only Action Recognition for Edge Deployment
- 融合层级结构与轻量化操作,仅用0.99万参数实现高精度动作识别。
- 在医疗退化条件下仍保持91%以上召回率,遮挡导致性能下降最显著。
- 适合资源受限的辅助监控场景,如智能养老或可穿戴设备部署。
资源受限的辅助监控需要紧凑的本地视频感知能力,并明确识别可靠性在常见视觉退化下的变化。本文提出一种紧凑的仅含RGB的动作识别网络,结合X3D式层级结构与时间移位、3D通用倒置瓶颈、Ghost逐点卷积、分解深度卷积算子、选择性时间适应及无参注意力机制。模型在NTU RGB+D 60(X-Sub/X-View)上达到95.10%/98.31%,在NTU RGB+D 120(X-Sub/X-Set)上达到90.88%/92.66%,参数仅0.96-0.99M。为连接紧凑部署与辅助应用,我们在未微调的情况下评估冻结检查点在官方医学条件组上的表现,保持原60-/120分类空间。医疗宏召回率在NTU60达93.19%/97.93%,在NTU120达91.47%/91.48%。在NTU60 X-Sub的可控二级视觉压力下,低光导致召回率下降1.77点,运动模糊下降2.23点,裁剪扰动下降4.87点,遮挡下降19.05点;四类扰动中遮挡影响最大。Jetson Orin Nano上TensorRT FP16引擎占用5.30 MiB,体现强静态紧凑性但吞吐非最优。结果表明该模型适合作为低频辅助监控中的紧凑感知组件,而非完整临床或机器人系统。
原文摘要 · Abstract (English)
Resource-constrained assistive monitoring requires compact local video perception and an explicit understanding of how recognition reliability changes under common visual degradation. We present a compact RGB-only action-recognition network that combines an X3D-style hierarchy with temporal shift, a 3D Universal Inverted Bottleneck, Ghost pointwise convolutions, factorized depthwise operators, selective temporal adaptation, and parameter-free attention. The model attains 95.10/98.31% on NTU RGB+D 60 (X-Sub/X-View) and 90.88/92.66% on NTU RGB+D 120 (X-Sub/X-Set) with only 0.96-0.99M parameters. To connect compact deployment with assistive use, we evaluate frozen checkpoints without retraining on the datasets' official Medical Conditions groups while retaining the original 60-/120-way decision spaces. Medical macro recall reaches 93.19/97.93% on NTU60 and 91.47/91.48% on NTU120. Under controlled Level-2 visual stress on NTU60 X-Sub, medical recall drops by 1.77 points under low light, 2.23 under motion blur, 4.87 under crop perturbation, and 19.05 under occlusion; among the four selected Level-2 perturbations, occlusion causes the largest observed degradation. On Jetson Orin Nano, the TensorRT FP16 engine occupies 5.30 MiB, demonstrating strong static compactness but not throughput superiority. The results position the model as a compact perception component for lower-rate assistive monitoring, rather than as a complete clinical or robotic system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。