arXiv:2602.22417cs.SDeess.AS2026-02中稿 · Interspeech 2026

用离散扩散模型提升语音降噪,尤其在低信噪比下效果显著。

Absorbing Discrete Diffusion for Speech Enhancement

  • 结合神经音频编码器与扩散模型,非自回归建模语音码本
  • 在两个数据集上达到领先指标,低信噪比下性能更优
  • 适合关注高效语音增强与非自回归生成的研究者

受神经语音编码和基于扩散的语言建模的启发,本文提出一种基于吸收式离散扩散的语音增强方法(ADDSE),通过建模给定含噪语音码本时干净语音码本的条件分布来实现。该方法利用神经音频编解码器的丰富潜在空间与扩散模型的非自回归采样特性。为高效建模残差向量量化码本的分层结构,提出RQDiT,融合RQ-Transformer与扩散Transformer的技术,实现非自回归建模。实验结果表明,在两个数据集上,该方法在非侵入性客观指标上表现优异,尤其在低信噪比和少采样步数条件下。代码与音频示例已公开。

原文摘要 · Abstract (English)

Inspired by recent developments in neural speech coding and diffusion-based language modeling, we tackle speech enhancement by modeling the conditional distribution of clean speech codes given noisy speech codes using absorbing discrete diffusion. The proposed approach, which we call ADDSE, leverages both the expressive latent space of neural audio codecs and the non-autoregressive sampling procedure of diffusion models. To efficiently model the hierarchical structure of residual vector quantization codes, we propose RQDiT, which combines techniques from RQ-Transformer and diffusion Transformers for non-autoregressive modeling. Results show competitive performance in terms of non-intrusive objective metrics on two datasets, especially at low signal-to-noise ratios and with few sampling steps. Code and audio examples are available online.

语音增强扩散模型非自回归码本建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。