用知识锚定与课程学习,让失语者语音合成更清晰自然。
Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning
- 通过师生模型+音频增强实现零样本多说话人语音合成
- 合成语音错误率显著降低,说话人特征保持度高
- 适合资源匮乏的失语者个性化语音生成研究
失语症患者因发音器官运动控制障碍导致语音可懂度大幅下降,难以采集足够长且清晰的语音数据用于训练个性化语音合成模型。针对这一问题,本文将任务视为领域迁移问题,提出基于知识锚定的师生模型框架,并结合音频增强的课程学习策略。实验表明,所提零样本多说话人语音合成模型能有效生成明显减少发音错误、保留高说话人相似度且韵律自然的合成语音,为失语者个性化语音重建提供了可行方案。
原文摘要 · Abstract (English)
Dysarthric speakers experience substantial communication challenges due to impaired motor control of the speech apparatus, which leads to reduced speech intelligibility. This creates significant obstacles in dataset curation since actual recording of long, articulate sentences for the objective of training personalized TTS models becomes infeasible. Thus, the limited availability of audio data, in addition to the articulation errors that are present within the audio, complicates personalized speech synthesis for target dysarthric speaker adaptation. To address this, we frame the issue as a domain transfer task and introduce a knowledge anchoring framework that leverages a teacher-student model, enhanced by curriculum learning through audio augmentation. Experimental results show that the proposed zero-shot multi-speaker TTS model effectively generates synthetic speech with markedly reduced articulation errors and high speaker fidelity, while maintaining prosodic naturalness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。