用向量量化提升语音增强模型鲁棒性,不依赖离散编码也能生效
Towards Robust Generative Speech Enhancement Using Vector Quantisation-Based Neural Audio Codec
- 在连续与离散潜在空间分别构建生成式语音增强框架
- 全微调的连续模型在DNS-MOS上超越所有离散模型变体
- 向量量化自带的纯净语音先验正则化是关键优势,可迁移至其他模型
本文研究基于向量量化(VQ)的神经音频编解码器(NAC)在语音增强(SE)中连续与离散潜在空间的建模策略,以及VQ正则化的作用。提出cNAC-SE和dNAC-SE两种框架,分别预测潜在空间中的连续表示与离散码本索引。通过理论分析与潜在空间可视化揭示其内在建模机制。实验结果表明,全微调的cNAC-SE模型在多种测试条件下均一致优于所有dNAC-SE变体,在DNS-MOS指标上达到现有生成式方法领先水平。与判别式模型对比显示,VQ通过纯净语音先验约束正则化实现鲁棒性提升,该效应独立于离散码本处理,凸显了VQ正则化对其他连续建模方法的可迁移价值。
原文摘要 · Abstract (English)
This work investigates modelling strategies in continuous and discrete latent spaces in the vector quantisation (VQ)-based neural audio codec (NAC) speech enhancement (SE), along with the role of VQ regularisation. We propose cNAC-SE and dNAC-SE frameworks that predict continuous representations and discrete tokens in latent space, respectively. Theoretical analysis and visualisations in latent space are performed to exhibit their inherent modelling mechanisms. Experimental results show that the fully fine-tuned cNAC-SE model consistently outperforms all dNAC-SE variants across diverse test conditions and achieves leading performance among established generative approaches in DNS-MOS metrics. Comparison with the discriminative counterpart shows that VQ enhances robustness through an intrinsic effect of clean-prior-constrained regularisation, independent of discrete token processing. This highlights the transferable value of VQ regularisation to other continuous modelling methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。