arXiv:2409.03377cs.SDcs.AI2024-09被引 4

aTENNuate用深度状态空间模型实现实时原始语音增强,效果优于现有方法。

aTENNuate: Optimized Real-time Speech Enhancement with Deep SSMs on Raw Audio

  • 基于深度状态空间自编码器,端到端处理原始波形,适合实时应用。
  • 在VoiceBank+DEMAND和DNS1数据集上PESQ得分更高,参数少、延迟低。
  • 可在4kHz、4bit低码率下保持良好性能,适合资源受限场景。

我们提出aTENNuate,一种针对原始语音增强的高效在线端到端深度状态空间自编码器。该网络主要在原始语音去噪任务上评估,还测试了超分辨率与去量化能力。在VoiceBank + DEMAND和Microsoft DNS1合成测试集上进行基准测试,aTENNuate在PESQ得分、参数量、计算量(MACs)和延迟方面均优于先前实时去噪模型。即使作为原始波形处理模型,仍能保持对干净信号的高保真度,且几乎无听觉失真。此外,当输入噪声语音被压缩至4000Hz和4比特时,模型仍表现良好,表明其在低资源环境下的通用语音增强能力。可通过pip install attenuate试用。

原文摘要 · Abstract (English)

We present aTENNuate, a simple deep state-space autoencoder configured for efficient online raw speech enhancement in an end-to-end fashion. The network's performance is primarily evaluated on raw speech denoising, with additional assessments on tasks such as super-resolution and de-quantization. We benchmark aTENNuate on the VoiceBank + DEMAND and the Microsoft DNS1 synthetic test sets. The network outperforms previous real-time denoising models in terms of PESQ score, parameter count, MACs, and latency. Even as a raw waveform processing model, the model maintains high fidelity to the clean signal with minimal audible artifacts. In addition, the model remains performant even when the noisy input is compressed down to 4000Hz and 4 bits, suggesting general speech enhancement capabilities in low-resource environments. Try it out by pip install attenuate

语音增强状态空间实时处理低码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。