arXiv:2412.13933eess.AScs.LG2024-12中稿 · ICASSP 2025 Satell…被引 1

用扩散模型增强帕金森失语语音会丢失关键病理特征

Investigating the Effects of Diffusion-based Conditional Generative Speech Models Used for Speech Enhancement on Dysarthric Speech

  • 用预训练扩散模型处理无噪失语语音,会误删病理声学特征
  • 增强后语音的失语线索显著减少,影响自动检测效果
  • 残留信号可补充失语特征,适合融合改进识别

本研究首次探索预训练条件生成语音模型对帕金森病失语语音的影响。采用基于扩散的语音增强模型,该模型此前在无噪正常语音上训练,学习清洁语音分布。实验发现,在理想无噪环境下处理失语语音时,部分未见过的非典型副语言特征被模型误去除。通过失语语音自动检测任务验证,增强过程会削弱关键声学特征。进一步表明,增强产生的残留语音信号在特征空间与原始语音融合后,可提供互补的失语线索。

原文摘要 · Abstract (English)

In this study, we aim to explore the effect of pre-trained conditional generative speech models for the first time on dysarthric speech due to Parkinson's disease recorded in an ideal/non-noisy condition. Considering one category of generative models, i.e., diffusion-based speech enhancement, these models are previously trained to learn the distribution of clean (i.e, recorded in a noise-free environment) typical speech signals. Therefore, we hypothesized that when being exposed to dysarthric speech they might remove the unseen atypical paralinguistic cues during the enhancement process. By considering the automatic dysarthric speech detection task, in this study, we experimentally show that during the enhancement process of dysarthric speech data recorded in an ideal non-noisy environment, some of the acoustic dysarthric speech cues are lost. Therefore such pre-trained models are not yet suitable in the context of dysarthric speech enhancement since they manipulate the pathological speech cues when they process clean dysarthric speech. Furthermore, we show that the removed acoustics cues by the enhancement models in the form of residue speech signal can provide complementary dysarthric cues when fused with the original input speech signal in the feature space.

语音增强失语症扩散模型病理语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。