arXiv:2510.09974cs.SD2025-10被引 4

将语音增强转为离散分类任务,统一处理多种噪声干扰。

Universal Discrete-Domain Speech Enhancement

  • 用预训练语音编解码器的离散码本,将语音增强转为逐层预测离散码字
  • 在含加性噪声、混响、削波等10种干扰组合下,语音质量提升2.5~3.8分(PESQ)
  • 适合需要高泛化能力的实用语音系统,尤其对复杂真实环境有效

真实场景中语音常受多种干扰影响,语音增强(SE)对鲁棒语音处理至关重要。然而现有方法多仅针对单一或有限类型失真(如加性噪声、混响、带宽限制),对多重同时失真的研究仍不足,制约了实际应用中的泛化能力。为此,本文提出一种新型通用离散域语音增强模型UDSE。不同于回归类方法直接预测连续波形或特征,UDSE将增强任务重新定义为离散域分类:利用预训练神经语音编解码器的残差向量量化(RVQ)码本,对清洁语音的离散码字进行预测。具体地,先提取受损语音的全局特征,再依序预测各层RVQ的码字,后一码字依赖前一码字结果。训练采用教师强制策略,以交叉熵损失优化。实验表明,该模型能有效应对加性噪声、混响、带宽限制、削波、相位失真、压缩失真及其组合等多种典型与非典型失真,在PESQ指标上平均提升2.5~3.8分,显著优于先进回归式方法,展现出更强的通用性与实用性。

原文摘要 · Abstract (English)

In real-world scenarios, speech signals are inevitably corrupted by various types of interference, making speech enhancement (SE) a critical task for robust speech processing. However, most existing SE methods only handle a limited range of distortions, such as additive noise, reverberation, or band limitation, while the study of SE under multiple simultaneous distortions remains limited. This gap affects the generalization and practical usability of SE methods in real-world environments.To address this gap, this paper proposes a novel Universal Discrete-domain SE model called UDSE.Unlike regression-based SE models that directly predict clean speech waveform or continuous features, UDSE redefines SE as a discrete-domain classification task, instead predicting the clean discrete tokens quantized by the residual vector quantizer (RVQ) of a pre-trained neural speech codec.Specifically, UDSE first extracts global features from the degraded speech. Guided by these global features, the clean token prediction for each VQ follows the rules of RVQ, where the prediction of each VQ relies on the results of the preceding ones. Finally, the predicted clean tokens from all VQs are decoded to reconstruct the clean speech waveform. During training, the UDSE model employs a teacher-forcing strategy, and is optimized with cross-entropy loss. Experimental results confirm that the proposed UDSE model can effectively enhance speech degraded by various conventional and unconventional distortions, e.g., additive noise, reverberation, band limitation, clipping, phase distortion, and compression distortion, as well as their combinations. These results demonstrate the superior universality and practicality of UDSE compared to advanced regression-based SE methods.

语音增强离散建模通用性编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。