统一评估六种基于音频编码器潜空间的生成式语音增强方法,提升音质评测指标。
Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec
- 构建覆盖六类建模范式的统一框架,使用离散/连续潜变量特征。
- 连续域方法整体优于离散域,非自回归模型表现最佳。
- 引入重构辅助损失,显著提升多种客观评价指标。
基于语言模型(LM)的语音增强(SE)近年来迅速发展,利用神经音频编码器(NAC)的潜空间特征。本文首先提出一个统一框架,涵盖六种主流的基于离散/连续潜变量的LM生成式SE建模范式:离散/连续自回归(D/CAR)、离散/连续非自回归(D/CNAR)、离散扩散(DDiff)与连续流匹配(CFM)。其次,首次在统一实验设置下,通过多样化的主观与非主观评价指标进行公平、全面的性能对比。第三,提出一种在重建语音上施加辅助损失的微调策略,有效提升各类指标。在URGENT 2025语音增强挑战赛数据集上训练与评估,所有连续域范式均优于其离散域对应方法,其中非自回归模型(CNAR)表现最佳。进一步验证表明,该辅助损失策略在全部六种范式中一致提升了DNSMOS、NISQA、PESQ与POLQA得分。
原文摘要 · Abstract (English)
Language model (LM)-based speech enhancement (SE) has recently emerged rapidly using latent space features of neural audio codecs (NACs). In this paper, first, we present a unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features: discrete or continuous autoregressive (D/CAR) SE, discrete or continuous non-autoregressive (D/CNAR) SE, discrete diffusion (DDiff) SE, and continuous flow matching (CFM) SE. Second, we are the first to compare their performance in a unified experimental setup and synopsis with diverse intrusive and non-intrusive metrics, enabling a fair and comprehensive evaluation. Third, we propose a fine-tuning strategy with auxiliary losses on reconstructed speech to improve both intrusive and non-intrusive metrics. Trained and evaluated on URGENT 2025 Speech Enhancement Challenge data splits, all continuous-domain paradigms excel their discrete-domain counterparts. The overall best approach turns out to be CNAR. We further show that our proposed auxiliary loss fine-tuning strategy helps to improve DNSMOS, NISQA, PESQ, and POLQA consistently in all six paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。