arXiv:2409.18239cs.SDcs.LG2024-09被引 6

用轻量模型实现亚毫秒级语音增强,提升听戴设备体验

Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables

  • 用轻量LSTM生成最小相位FIR滤波器系数,支持逐样本处理
  • 实测平均算法延迟0.32~1.25毫秒,单麦克风下SI-SDRi达4.1 dB
  • 可在低功耗芯片上运行,适合资源受限的听戴设备

低延迟模型对助听器和听戴设备等实时语音增强应用至关重要。然而,针对资源受限听戴设备的亚毫秒级延迟方案仍研究不足。本文展示一种计算高效的最小相位FIR滤波器方法,实现逐样本处理,平均算法延迟为0.32至1.25毫秒。在单麦克风条件下,平均SI-SDRi达到4.1 dB。该方法在未见音频上表现出良好泛化能力,DNSMOS提升0.2。采用仅626k参数的轻量LSTM模型生成FIR系数。在低功耗DSP上的真实硬件实现中,系统运行仅需376 MIPS,平均端到端延迟为3.35毫秒。此外,本文还与现有低延迟谱掩蔽技术进行了对比。希望本工作能深化对延迟的理解,助力提升听戴设备的舒适性与可用性。

原文摘要 · Abstract (English)

Low latency models are critical for real-time speech enhancement applications, such as hearing aids and hearables. However, the sub-millisecond latency space for resource-constrained hearables remains underexplored. We demonstrate speech enhancement using a computationally efficient minimum-phase FIR filter, enabling sample-by-sample processing to achieve mean algorithmic latency of 0.32 ms to 1.25 ms. With a single microphone, we observe a mean SI-SDRi of 4.1 dB. The approach shows generalization with a DNSMOS increase of 0.2 on unseen audio recordings. We use a lightweight LSTM-based model of 626k parameters to generate FIR taps. Using a real hardware implementation on a low-power DSP, our system can run with 376 MIPS and a mean end-to-end latency of 3.35 ms. In addition, we provide a comparison with existing low-latency spectral masking techniques. We hope this work will enable a better understanding of latency and can be used to improve the comfort and usability of hearables.

语音增强低延迟听戴设备FIR滤波器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。