用双路编码融合确定性增强与噪声特征,提升扩散模型语音增强效果
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
- 设计双流编码结构,同时利用确定性增强和原始噪声特征
- 在CHiME4数据集上达到更优的客观指标与更稳定性能
- 结合粗粒度与细粒度处理,兼顾评估得分与生成稳定性
基于扩散模型的语音增强需引入可靠先验条件以生成准确结果。然而,直接使用噪声特征作为条件存在可靠性问题。一种方案是采用确定性方法增强后的特征作为条件,但该过程可能导致信息失真或损失,影响扩散过程。本文首先研究不同确定性语音增强模型作为条件的效果,对比仅使用增强特征(仅确定性)与同时使用增强与噪声特征(确定性-噪声)两种方式。初步实验表明,使用确定性增强特征可改善真实场景下的听感体验,而具体选择取决于所用确定性模型。基于此,提出修复-扩散模型(DERDM-SE),通过双流编码有效融合两类条件。进一步发现,细粒度确定性模型在客观评价指标上潜力更大,而基于UNet的模型提供更稳定的扩散性能。因此,DERDM-SE设计了一种结合粗粒度与细粒度处理的确定性模块。在CHiME4数据集上的实验显示,该模型显著提升了语音增强评分,并表现出优于其他扩散模型的稳定性。
原文摘要 · Abstract (English)
Diffusion-based speech enhancement (SE) models need to incorporate correct prior knowledge as reliable conditions to generate accurate predictions. However, providing reliable conditions using noisy features is challenging. One solution is to use features enhanced by deterministic methods as conditions. However, the information distortion and loss caused by deterministic methods might affect the diffusion process. In this paper, we first investigate the effects of using different deterministic SE models as conditions for diffusion. We validate two conditions depending on whether the noisy feature was used as part of the condition: one using only the deterministic feature (deterministic-only), and the other using both deterministic and noisy features (deterministic-noisy). Preliminary investigation found that using deterministic enhanced conditions improves hearing experiences on real data, while the choice between using deterministic-only or deterministic-noisy conditions depends on the deterministic models. Based on these findings, we propose a dual-streaming encoding Repair-Diffusion Model for SE (DERDM-SE) to more effectively utilize both conditions. Moreover, we found that fine-grained deterministic models have greater potential in objective evaluation metrics, while UNet-based deterministic models provide more stable diffusion performance. Therefore, in the DERDM-SE, we propose a deterministic model that combines coarse- and fine-grained processing. Experimental results on CHiME4 show that the proposed models effectively leverage deterministic models to achieve better SE evaluation scores, along with more stable performance compared to other diffusion-based SE models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。