用对比学习提升失语症语音重建质量,保留说话人特征。
Enhancement of Dysarthric Speech Reconstruction by Contrastive Learning
- 用对比学习提取说话人嵌入,结合XLS-R特征替代滤波器组。
- 重建语音在自然度、可懂度和说话人一致性上显著提升。
- 特别适合女性失语症患者,对中度到中重度患者效果明显。
由于失语症语音存在病理声学模式,其重建极具挑战性,尤其在缺乏正常语音作为参考时,保持说话人身份尤为困难。本文提出一种新方法,利用对比学习提取说话人嵌入用于重建,并采用XLS-R表示代替传统滤波器组。实验结果显示,重建语音在语音质量、自然度、可懂度、说话人身份保持及性别一致性方面均有提升。使用Jasper语音识别系统评估,中度和中重度失语症患者的平均意见分(MOS)分别提升1.51和2.12,词错误率分别降低25.45%和32.1%。该方法为失语症语音重建提供了有效解决方案。
原文摘要 · Abstract (English)
Dysarthric speech reconstruction is challenging due to its pathological sound patterns. Preserving speaker identity, especially without access to normal speech, is a key challenge. Our proposed approach uses contrastive learning to extract speaker embedding for reconstruction, while employing XLS-R representations instead of filter banks. The results show improved speech quality, naturalness, intelligibility, speaker identity preservation, and gender consistency for female speakers. Reconstructed speech exhibits 1.51 and 2.12 MOS score improvements and reduces word error rates by 25.45% and 32.1% for moderate and moderate-severe dysarthria speakers using Jasper speech recognition system, respectively. This approach offers promising advancements in dysarthric speech reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。