用扩散模型提升失语症语音可懂度,改善自动识别效果。
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
- 采用扩散模型与传统信号处理法增强失语症语音。
- 增强后语音在客观指标上更接近正常语音分布。
- 适合语音识别研究者及康复技术开发者参考。
失语症语音因高度变异性和低可懂度,给自动语音识别(ASR)系统带来挑战。本文探索使用扩散模型进行失语症语音增强,假设该方法能使失语症语音分布更接近正常语音,从而提升识别性能。我们评估了两种基于扩散模型和一种基于信号处理的语音增强算法,在两个英文失语症语音语料库上的表现,包括语音可懂度与质量。对正常与失语症语音均应用增强,并使用Whisper-Turbo评估增强前后语音的ASR性能;同时通过主观与客观方式评估原始与增强后失语症语音的质量。此外,还对Whisper-Turbo在增强语音上进行微调,以分析其对识别性能的影响。
原文摘要 · Abstract (English)
Dysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement, which is based on the hypothesis that using diffusion-based speech enhancement moves the distribution of dysarthric speech closer to that of typical speech, which could potentially improve dysarthric speech recognition performance. We assess the effect of two diffusion-based and one signal-processing-based speech enhancement algorithms on intelligibility and speech quality of two English dysarthric speech corpora. We applied speech enhancement to both typical and dysarthric speech and evaluate the ASR performance using Whisper-Turbo, and the subjective and objective speech quality of the original and enhanced dysarthric speech. We also fine-tuned Whisper-Turbo on the enhanced speech to assess its impact on recognition performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。