通过双路径结构提升DeepFilterNet2的语音增强效果,兼顾实时性与低功耗部署。
DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN
- 采用双路径RNN增强长时依赖和跨频带建模能力。
- 在VoiceBank+DEMAND和DNS4上优于DeepFilterNet2,多语言低信噪比测试中表现领先。
- 支持边缘设备实时运行,适合始终在线的嵌入式应用。
本文提出DPDFNet,一种因果单通道语音增强模型,通过在编码器中引入双路径块,强化了长时依赖与跨频带建模能力,同时保留原始增强框架。此外,我们证明添加抑制过度衰减的损失项,并配合针对“始终在线”应用设计的微调阶段,可显著提升整体性能。在标准VoiceBank+DEMAND和DNS4盲测基准上,DPDFNet持续优于DeepFilterNet2,且在与其他因果开源模型对比中表现强劲。我们还构建了一个补充性的多语言低信噪比评估集,包含12种语言的长录音,涵盖日常噪声场景,结果显示DPDFNet在该数据集上优于其他因果开源模型,包括一些更大、计算更密集的模型。我们提出综合指标PRISM,为侵入式与非侵入式指标的归一化聚合,其性能随双路径块数量提升而明显改善。此外,我们在Ceva-NeuPro-Nano边缘NPUs上验证了模型的在设备可行性:DPDFNet-4(第二大的模型)在NPN32上实现实时性能,在NPN64上运行更快,证实了在严格功耗与延迟约束下仍可保持顶尖音质。
原文摘要 · Abstract (English)
We present DPDFNet, a causal single-channel speech enhancement model that extends DeepFilterNet2 architecture with dual-path blocks in the encoder, strengthening long-range temporal and cross-band modeling while preserving the original enhancement framework. In addition, we demonstrate that adding a loss component to mitigate over-attenuation in the enhanced speech, combined with a fine-tuning phase tailored for "always-on" applications, leads to substantial improvements in overall model performance. We evaluate DPDFNet on the standard VoiceBank+DEMAND and DNS4 blind test benchmarks, where it shows consistent gains over DeepFilterNet2 and strong overall performance against other causal open-source models. In addition, we introduce a supplementary multilingual low-SNR evaluation set comprising long recordings in 12 languages across everyday noise scenarios, on which DPDFNet delivers superior performance to other causal open-source models, including some that are substantially larger and more computationally demanding. We also propose an holistic metric named PRISM, a composite, scale-normalized aggregate of intrusive and non-intrusive metrics, which demonstrates clear scalability with the number of dual-path blocks. We further demonstrate on-device feasibility by deploying DPDFNet on Ceva-NeuPro-Nano edge NPUs. Results indicate that DPDFNet-4, our second-largest model, achieves real-time performance on NPN32 and runs even faster on NPN64, confirming that state-of-the-art quality can be sustained within strict embedded power and latency constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。