轻量级手势识别框架,实时速度提升13倍
Duo Streamers: A Streaming Gesture Recognition Framework
- 三阶段稀疏识别+轻量RNN模型,降低计算开销
- 实时因子降低92.3%,参数量仅为主流模型1/38~1/9
- 适合边缘设备部署,适用于多模态场景
资源受限场景下的手势识别面临高精度与低延迟难以兼顾的挑战。本文提出流式手势识别框架Duo Streamers,通过三阶段稀疏识别机制、带有外部隐状态的RNN-lite模型,以及专用训练与后处理流程,在实时性能和轻量化设计上取得创新进展。实验表明,Duo Streamers在准确率上媲美主流方法,同时将实时因子降低约92.3%,实现近13倍加速;参数量相比主流模型减少至1/38(空闲状态)和1/9(忙碌状态)。该框架不仅为资源受限设备提供了高效实用的手势识别方案,也为多模态和多样化应用场景奠定了基础。
原文摘要 · Abstract (English)
Gesture recognition in resource-constrained scenarios faces significant challenges in achieving high accuracy and low latency. The streaming gesture recognition framework, Duo Streamers, proposed in this paper, addresses these challenges through a three-stage sparse recognition mechanism, an RNN-lite model with an external hidden state, and specialized training and post-processing pipelines, thereby making innovative progress in real-time performance and lightweight design. Experimental results show that Duo Streamers matches mainstream methods in accuracy metrics, while reducing the real-time factor by approximately 92.3%, i.e., delivering a nearly 13-fold speedup. In addition, the framework shrinks parameter counts to 1/38 (idle state) and 1/9 (busy state) compared to mainstream models. In summary, Duo Streamers not only offers an efficient and practical solution for streaming gesture recognition in resource-constrained devices but also lays a solid foundation for extended applications in multimodal and diverse scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。