arXiv:2601.21940eess.AS2026-01中稿 · IEEE ICASSP 2026被引 3

提出单步扩散语音增强模型,兼顾音质与语音准确性。

DisContSE: Single-Step Diffusion Speech Enhancement Based on Joint Discrete and Continuous Embeddings

  • 联合离散编码令牌与连续嵌入,分模块提升音质与可懂度。
  • 单步推理实现高效增强,在多项指标上超越基线方法。
  • 适合追求高保真与高语音准确性的语音增强研究者使用。

基于离散音频编码器特征的扩散模型在语音增强领域备受关注,因其具备更强的语音成分重建能力。然而,这类方法通常因多轮逆向过程导致推理计算复杂度高,且在非侵入式指标上表现优异,但在侵入式指标上表现不佳,难以准确重构音素。本文提出DisContSE,一种基于离散编码令牌与连续嵌入联合建模的高效扩散语音增强模型。首先,分别设计离散与连续增强模块,分别作用于离散编码令牌和连续嵌入,以同时提升音质与可懂度;其次,引入语义增强模块以优化音素准确性;第三,通过新颖的量化误差掩码初始化策略,实现单步推理,据我们所知,这是首个基于音频编码器的单步扩散语音增强方法。在URGENT 2024语音增强挑战赛数据集上训练与评估,DisContSE在PESQ、POLQA、UTMOS以及主观ITU-T P.808听感测试中均优于报告的时频域扩散基线方法,整体表现位列第一。

原文摘要 · Abstract (English)

Diffusion speech enhancement on discrete audio codec features gain immense attention due to their improved speech component reconstruction capability. However, they usually suffer from high inference computational complexity due to multiple reverse process iterations. Furthermore, they generally achieve promising results on non-intrusive metrics but show poor performance on intrusive metrics, as they may struggle in reconstructing the correct phones. In this paper, we propose DisContSE, an efficient diffusion-based speech enhancement model on joint discrete codec tokens and continuous embeddings. Our contributions are three-fold. First, we formulate both a discrete and a continuous enhancement module operating on discrete audio codec tokens and continuous embeddings, respectively, to achieve improved fidelity and intelligibility simultaneously. Second, a semantic enhancement module is further adopted to achieve optimal phonetic accuracy. Third, we achieve a single-step efficient reverse process in inference with a novel quantization error mask initialization strategy, which, according to our knowledge, is the first successful single-step diffusion speech enhancement based on an audio codec. Trained and evaluated on URGENT 2024 Speech Enhancement Challenge data splits, the proposed DisContSE excels top-reported time- and frequency-domain diffusion baseline methods in PESQ, POLQA, UTMOS, and in a subjective ITU-T P.808 listening test, clearly achieving an overall top rank.

语音增强扩散模型单步推理离散编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。