用多视角判别增强扩散模型,提升语音增强效果与效率
MDDM: A Multi-view Discriminative Enhanced Diffusion-based Model for Speech Enhancement
- 融合时域、频域、噪声三维度特征进行判别预测
- 仅需少量采样步骤即可达到先进水平的语音质量
- 适合追求高效高保真语音增强的研究与应用
随着深度学习的发展,语音增强在语音质量方面已取得显著优化。以往方法通常依赖判别式监督学习或生成建模,往往导致语音失真或计算成本过高。本文提出MDDM——一种多视角判别增强的扩散模型。具体而言,将时域、频域和噪声三个领域的特征作为判别预测网络的输入,生成初步频谱图;随后通过若干次推理采样步骤,将判别输出转换为干净语音。由于判别输出与目标干净语音分布存在重叠,因此只需较少采样步数即可达到与其他扩散模型相当的性能。在公开数据集和真实场景数据集上的实验验证了MDDM的有效性,无论在主观评价还是客观指标上均表现优异。
原文摘要 · Abstract (English)
With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which tends to introduce speech distortions or high computational cost. In this paper, we propose MDDM, a Multi-view Discriminative enhanced Diffusion-based Model. Specifically, we take the features of three domains (time, frequency and noise) as inputs of a discriminative prediction network, generating the preliminary spectrogram. Then, the discriminative output can be converted to clean speech by several inference sampling steps. Due to the intersection of the distributions between discriminative output and clean target, the smaller sampling steps can achieve the competitive performance compared to other diffusion-based methods. Experiments conducted on a public dataset and a realworld dataset validate the effectiveness of MDDM, either on subjective or objective metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。