轻量模型在嵌入式芯片上实现低延迟语音增强,逼近助听器部署门槛。
Feasibility of Time-Domain DNN-Based Speech Enhancement on Embedded FPGA for Hearing Aids

- 用轻量级SuDoRM-RF++模型在Kria KV260上实测
- 定点16位下首样本延迟仅9.7毫秒,达标临床要求
- 精度降低一半内存减半,适合资源受限设备
助听器对延迟和功耗有严格限制,当前基于深度神经网络(DNN)的语音增强系统难以在嵌入式硬件上满足要求。本文通过在AMD-Xilinx Kria KV260上部署轻量级SuDoRM-RF++架构进行语音分离与降噪,分别在FP32和16位定点精度下评估。结果显示,首样本延迟受片上参数缓存影响,而非计算吞吐量,表明数据移动是主要瓶颈。精度降低使模型内存占用减半,且未牺牲客观语音质量。定点降噪加速器实现9.7~16.0毫秒的首样本延迟,满足10毫秒临床阈值;语音分离达16.0毫秒。这些数据明确了嵌入式DNN语音增强的资源需求,并量化了与助听器实际部署间的差距。
原文摘要 · Abstract (English)
Hearing aids impose strict latency and power constraints that current DNN-based speech enhancement systems struggle to meet on embedded hardware. We characterize this gap by deploying both speech separation and denoising using the lightweight SuDoRM-RF++ architecture on the AMD-Xilinx Kria KV260, evaluated at FP32 and 16-bit fixed-point precision for each task. Across these configurations, first-sample latency tracks with on-chip parameter caching rather than arithmetic throughput, identifying data movement as the primary bottleneck. Precision reduction halves the model memory footprint without compromising objective speech quality. The fixed-point denoising accelerator achieves a first-sample latency of 9.7~ms, meeting the 10~ms clinical threshold, while speech separation reaches 16.0~ms. These measurements establish concrete resource requirements for embedded DNN-based speech enhancement and quantify the remaining gap to hearing aid deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。