系统梳理7类高效模型技术,助力语音关键词识别在低功耗设备落地
Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms
- 从模型架构到神经网络搜索,覆盖7类适配边缘设备的优化方法
- 提出适用于低功耗场景的轻量级关键词识别框架,支持实时运行
- 适合从事边缘智能、语音交互与TinyML研发的研究者参考
小尺寸关键词检测(SF-KWS)在智能语音设备、智能手机及物联网应用中日益重要,得益于深度学习的发展,可从连续语音流中识别预定义关键词。为在资源受限的边缘设备上实现低功耗、低内存的实时运行,亟需高效的微型机器学习(TinyML)框架。本文系统梳理了七类关键技术:模型架构、学习策略、模型压缩、注意力感知架构、特征优化、神经网络搜索及混合方法,均适用于构建高效SF-KWS系统。该综述为理解、应用或贡献于该领域提供重要参考,并揭示了来自自动语音识别及特定语音关键词检测领域的多个潜在研究方向。
原文摘要 · Abstract (English)
Small-Footprint Keyword Spotting (SF-KWS) has gained popularity in today's landscape of smart voice-activated devices, smartphones, and Internet of Things (IoT) applications. This surge is attributed to the advancements in Deep Learning, enabling the identification of predefined words or keywords from a continuous stream of words. To implement the SF-KWS model on edge devices with low power and limited memory in real-world scenarios, a efficient Tiny Machine Learning (TinyML) framework is essential. In this study, we explore seven distinct categories of techniques namely, Model Architecture, Learning Techniques, Model Compression, Attention Awareness Architecture, Feature Optimization, Neural Network Search, and Hybrid Approaches, which are suitable for developing an SF-KWS system. This comprehensive overview will serve as a valuable resource for those looking to understand, utilize, or contribute to the field of SF-KWS. The analysis conducted in this work enables the identification of numerous potential research directions, encompassing insights from automatic speech recognition research and those specifically pertinent to the realm of spoken SF-KWS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。