轻量级语音增强模型,适合移动端实时降噪。
HiFi-Stream: Streaming Speech Enhancement with Generative Adversarial Networks
- 基于HiFi++改进,压缩模型体积并降低计算开销。
- 在流式场景下性能优于现有主流方法。
- 适合资源受限设备部署,兼顾速度与效果。
语音增强技术已成为移动设备和语音软件的核心。然而,现代深度学习方案通常需要大量计算资源,难以在低资源设备上应用。本文提出HiFi-Stream,是近期发布的HiFi++模型的优化版本。实验表明,尽管模型尺寸和计算复杂度显著降低,其性能仍保持接近原始模型水平,成为目前最小、最快之一的语音增强模型。该模型在流式设置下进行评估,表现优于当前主流基线方法。
原文摘要 · Abstract (English)
Speech Enhancement techniques have become core technologies in mobile devices and voice software. Still, modern deep learning solutions often require high amount of computational resources what makes their usage on low-resource devices challenging. We present HiFi-Stream, an optimized version of recently published HiFi++ model. Our experiments demonstrate that HiFi-Stream saves most of the qualities of the original model despite its size and computational complexity improved in comparison to the original HiFi++ making it one of the smallest and fastest models available. The model is evaluated in streaming setting where it demonstrates its superior performance in comparison to modern baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。