针对帕金森病语音增强,数据增强有效提升效果但仍有差距。
Data Augmentation for Pathological Speech Enhancement
- 测试三类数据增强:变换、生成和噪声增强
- 噪声增强效果最好,生成增强过多反而降低性能
- 对预测型模型更有效,需针对性设计病理语音增强策略
当前先进语音增强(SE)模型在帕金森病患者语音上性能显著下降,主要因声学特征异常且数据稀缺。本文系统研究数据增强(DA)策略以改善此类语音的增强效果,评估了预测型与生成型SE模型。考察三类增强方法:变换类、生成类与噪声增强类,并通过客观指标分析其影响。实验表明,噪声增强带来最大且最稳定的性能提升;变换增强有中等改善;生成增强效果有限,且随合成数据量增加反而损害性能。此外,增强效果受模型类型影响,对预测型模型更为有益。尽管数据增强提升了病理语音的增强表现,但与正常语音之间仍存在性能差距,凸显未来需发展针对性的病理语音增强策略。
原文摘要 · Abstract (English)
The performance of state-of-the-art speech enhancement (SE) models considerably degrades for pathological speech due to atypical acoustic characteristics and limited data availability. This paper systematically investigates data augmentation (DA) strategies to improve SE performance for pathological speakers affected by Parkinson`s disease, evaluating both predictive and generative SE models. We examine three DA categories, i.e., transformative, generative, and noise augmentation, assessing their impact with objective SE metrics. Experimental results show that noise augmentation consistently delivers the largest and most robust gains, transformative augmentations provide moderate improvements, while generative augmentation yields limited benefits and can harm performance as the amount of synthetic data increases. Furthermore, we show that the effectiveness of DA varies depending on the SE model, with DA being more beneficial for predictive SE models. While our results demonstrate that DA improves SE performance for pathological speakers, a performance gap between neurotypical and pathological speech persists, highlighting the need for future research on targeted DA strategies for pathological speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。