arXiv:2608.12082eess.AS2026-08中稿 · IWAENC 2026

统一评估六种基于音频编码器潜空间的生成式语音增强方法,提升音质评测指标。

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec

  • 构建覆盖六类建模范式的统一框架,使用离散/连续潜变量特征。
  • 连续域方法整体优于离散域,非自回归模型表现最佳。
  • 引入重构辅助损失,显著提升多种客观评价指标。

基于语言模型(LM)的语音增强(SE)近年来迅速发展,利用神经音频编码器(NAC)的潜空间特征。本文首先提出一个统一框架,涵盖六种主流的基于离散/连续潜变量的LM生成式SE建模范式:离散/连续自回归(D/CAR)、离散/连续非自回归(D/CNAR)、离散扩散(DDiff)与连续流匹配(CFM)。其次,首次在统一实验设置下,通过多样化的主观与非主观评价指标进行公平、全面的性能对比。第三,提出一种在重建语音上施加辅助损失的微调策略,有效提升各类指标。在URGENT 2025语音增强挑战赛数据集上训练与评估,所有连续域范式均优于其离散域对应方法,其中非自回归模型(CNAR)表现最佳。进一步验证表明,该辅助损失策略在全部六种范式中一致提升了DNSMOS、NISQA、PESQ与POLQA得分。

原文摘要 · Abstract (English)

Language model (LM)-based speech enhancement (SE) has recently emerged rapidly using latent space features of neural audio codecs (NACs). In this paper, first, we present a unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features: discrete or continuous autoregressive (D/CAR) SE, discrete or continuous non-autoregressive (D/CNAR) SE, discrete diffusion (DDiff) SE, and continuous flow matching (CFM) SE. Second, we are the first to compare their performance in a unified experimental setup and synopsis with diverse intrusive and non-intrusive metrics, enabling a fair and comprehensive evaluation. Third, we propose a fine-tuning strategy with auxiliary losses on reconstructed speech to improve both intrusive and non-intrusive metrics. Trained and evaluated on URGENT 2025 Speech Enhancement Challenge data splits, all continuous-domain paradigms excel their discrete-domain counterparts. The overall best approach turns out to be CNAR. We further show that our proposed auxiliary loss fine-tuning strategy helps to improve DNSMOS, NISQA, PESQ, and POLQA consistently in all six paradigms.

语音增强潜空间生成模型音频编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。