μNet在嵌入式设备上实现极低内存与计算开销的语音增强。
μNet: Ultra-Low-Memory and Low-Complexity Speech Enhancement for Embedded Digital Signal Processors

- 设计端到端轻量级神经网络,仅需90KB静态内存和28MMACs
- 4ms延迟下性能媲美同类先进方法,支持整数运算
- 适配消费级DSP芯片,可部署于真实嵌入式场景
嵌入式数字信号处理器(DSP)上的语音增强对内存占用、计算复杂度、延迟及整数运算支持有严格限制。尽管近期基于深度神经网络(DNN)的方法分别缓解了其中部分挑战,但文献中尚无统一框架能同时满足所有要求以实现实际部署。本文提出μNet,一种超低内存、低复杂度、低延迟的端到端DNN模型。该方法仅需90 KB静态内存与28 MMACs,算法延迟低至4 ms,且性能可与同复杂度的最先进方法相当。实验表明,μNet兼容神经加速器,并可在消费级DSP平台(如Cadence Tensilica HiFi 4/5)上实现全整数运算。
原文摘要 · Abstract (English)
Speech enhancement on embedded digital signal processors (DSPs) imposes strict constraints on memory footprint, computational complexity, latency, and support for integer operations. Although recent DNN-based approaches have addressed these challenges individually, no unified framework in the literature simultaneously addresses all these requirements for practical deployment. In this work, we propose μNet, an ultra-low-memory, low-complexity, and low-latency end-to-end DNN model. The proposed method requires only $90$~KB of static memory and $28$~MMACs, while supporting an algorithmic latency as low as $4$~ms with performance comparable to state-of-the-art methods of similar complexity. Our experiments demonstrate that μNet is compatible with neural accelerators and supports full integer-arithmetic operations on consumer DSP platforms such as Cadence Tensilica HiFi 4/5.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。