arXiv:2602.01793cs.SD2026-02中稿 · ICASSP 2026

并行生成语音增强模型,提升效率与音质。

ParaGSE: Parallel Generative Speech Enhancement with Group-Vector-Quantization-based Neural Speech Codec

  • 采用分组向量量化编码器,实现并行令牌预测
  • 在多种噪声混叠下均优于传统方法,音质更佳
  • 支持高效并行计算,生成速度提升1.5倍

近期生成式语音增强受到广泛关注,但现有方法受限于复杂度高、效率低及音质不佳。本文提出一种新型并行生成语音增强(ParaGSE)框架,基于分组向量量化(GVQ)的神经语音编解码器,通过独立向量量化生成互不依赖的令牌,实现并行令牌预测。具体而言,ParaGSE利用该编解码器将受损语音编码为不同令牌,通过条件于受损频谱特征的并行分支预测对应干净令牌,并由编解码器解码重建清晰语音。实验表明,无论在噪声、混响、带宽限制及其混合等多种失真条件下,ParaGSE均持续优于判别式与生成式基线。此外,得益于令牌预测中的并行计算,其在CPU上的生成效率相较串行生成方法提升约1.5倍。

原文摘要 · Abstract (English)

Recently, generative speech enhancement has garnered considerable interest; however, existing approaches are hindered by excessive complexity, limited efficiency, and suboptimal speech quality. To overcome these challenges, this paper proposes a novel parallel generative speech enhancement (ParaGSE) framework that leverages a group vector quantization (GVQ)-based neural speech codec. The GVQ-based codec adopts separate VQs to produce mutually independent tokens, enabling efficient parallel token prediction in ParaGSE. Specifically, ParaGSE leverages the GVQ-based codec to encode degraded speech into distinct tokens, predicts the corresponding clean tokens through parallel branches conditioned on degraded spectral features, and ultimately reconstructs clean speech via the codec decoder. Experimental results demonstrate that ParaGSE consistently produces superior enhanced speech compared to both discriminative and generative baselines, under a wide range of distortions including noise, reverberation, band-limiting, and their mixtures. Furthermore, empowered by parallel computation in token prediction, ParaGSE attains about a 1.5-fold improvement in generation efficiency on CPU compared with serial generative speech enhancement approaches.

语音增强生成模型并行计算向量量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。