基于流匹配的语音增强模型,可同时应对多种噪声和失真。
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
- 在隐空间中使用流匹配与扩散Transformer结合,提升鲁棒性。
- 在五种未见失真上性能超越当前最优模型,参数仅用4.9%。
- 适合需要高适应性的实际语音增强场景,如智能设备、会议系统。
生成模型如扩散模型和流匹配在音频任务中表现优异。然而,语音增强(SE)模型通常在有限数据集上训练且评估条件狭窄,限制了实际应用。为此,我们提出DiT-Flow,一种基于流匹配的语音增强框架,依托隐空间中的扩散Transformer(DiT)主干网络,可抵御多种失真,包括噪声、混响和压缩。DiT-Flow在由LibriSpeech、FSD50K、FMA和90个Matterport3D场景构成的合成但声学真实的StillSonicSet数据集上验证。实验表明,该方法持续优于现有生成式语音增强模型,证明了流匹配在多条件语音增强中的有效性。尽管合成数据真实度不断提升,语音增强仍面临训练与部署条件不匹配的瓶颈。通过将LoRA与MoE框架结合,仅用总参数量的4.9%即可实现对五种未见失真的高性能鲁棒训练。
原文摘要 · Abstract (English)
Recent advances in generative models, such as diffusion and flow matching, have shown strong performance in audio tasks. However, speech enhancement (SE) models are typically trained on limited datasets and evaluated under narrow conditions, limiting real-world applicability. To address this, we propose DiT-Flow, a flow matching-based SE framework built on the latent Diffusion Transformer (DiT) backbone and trained for robustness across diverse distortions, including noise, reverberation, and compression. DiT-Flow operates on compact variational auto-encoders (VAEs)-derived latent features. We validated our approach on StillSonicSet, a synthetic yet acoustically realistic dataset composed of LibriSpeech, FSD50K, FMA, and 90 Matterport3D scenes. Experiments show that DiT-Flow consistently outperforms state-of-the-art generative SE models, demonstrating the effectiveness of flow matching in multi-condition speech enhancement. Despite ongoing efforts to expand synthetic data realism, a persistent bottleneck in SE is the inevitable mismatch between training and deployment conditions. By integrating LoRA with the MoE framework, we achieve both parameter-efficient and high-performance training for DiT-Flow robust to multiple distortions with using 4.9% percentage of the total parameters to obtain a better performance on five unseen distortions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。