对比生成与判别模型在降噪中的表现,评估其鲁棒性与实用价值。
A Comparison of Generative and Discriminative Methods for Speech Enhancement: Robustness, Complexity, and Hallucination

- 对比生成与判别模型在高低信噪比下的降噪能力
- 发现生成模型在低信噪比下更优但存在幻觉问题
- 适合关注实际部署成本与效果平衡的研究者
本研究对基于深度学习的生成与判别型语音增强方法进行了全面比较,重点关注噪声抑制任务。分析涵盖高、低信噪比条件下的性能表现,以及匹配与不匹配训练场景。考察了训练数据量、模型收敛速度对结果的影响,并从客观指标角度解析不同训练范式的表现差异。进一步比较了复杂度与性能之间的权衡及实际可行性。为强化评估,还研究了生成方法在词错误率与音素相似性上的幻觉特性。研究结果为科研人员和实践者判断各类方法的感知增益是否值得其计算开销提供了实证依据。
原文摘要 · Abstract (English)
In this study, we conduct a comprehensive comparative analysis of generative and discriminative deep learning-based speech enhancement methods, specifically in noise reduction tasks. Our investigation focuses on evaluating their effectiveness under high and low signal-to-noise ratio conditions, considering both matched and mismatched training scenarios. We further investigate the impact of training data volume, model convergence speed, and interpret the performance differences in terms of objective results for the considered training paradigms. Additionally, we compare the complexity-performance trade-off and the practical viability of these approaches. To further strengthen the evaluation, we study the hallucination characteristics of generative approaches in terms of word error rate and phoneme similarity. The insights derived from this study provide empirical evidence to assist researchers and practitioners in understanding whether the perceptual gains of different approaches justify their computational cost in practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。