通过原型解耦实现可控制的构音障碍语音合成
Prototype-Based Disentanglement for Controllable Dysarthric Speech Synthesis
- 用原型码本分离说话人音色与病理发音特征
- 在TORGO数据集上实现健康与障碍语音双向转换
- 适合语音康复和个性化辅助技术研究者
构音障碍语音具有高度变异性和标注数据稀缺,给自动语音识别(ASR)和辅助语音技术带来挑战。现有方法依赖合成数据增强或语音重建,但常将说话人身份与病理发音混杂,限制可控性与鲁棒性。本文提出ProtoDisent-TTS,基于预训练文本到语音模型,在统一潜在空间中解耦说话人音色与构音障碍特征。通过病理原型码本提供可解释、可调控的健康与障碍语音表示,结合带梯度反转层的双分类器目标,强制说话人嵌入对病理属性不变。在TORGO数据集上的实验表明,该设计可实现健康与障碍语音间的双向转换,带来一致的ASR性能提升,并实现鲁棒、说话人感知的语音重建。
原文摘要 · Abstract (English)
Dysarthric speech exhibits high variability and limited labeled data, posing major challenges for both automatic speech recognition (ASR) and assistive speech technologies. Existing approaches rely on synthetic data augmentation or speech reconstruction, yet often entangle speaker identity with pathological articulation, limiting controllability and robustness. In this paper, we propose ProtoDisent-TTS, a prototype-based disentanglement TTS framework built on a pre-trained text-to-speech backbone that factorizes speaker timbre and dysarthric articulation within a unified latent space. A pathology prototype codebook provides interpretable and controllable representations of healthy and dysarthric speech patterns, while a dual-classifier objective with a gradient reversal layer enforces invariance of speaker embeddings to pathological attributes. Experiments on the TORGO dataset demonstrate that this design enables bidirectional transformation between healthy and dysarthric speech, leading to consistent ASR performance gains and robust, speaker-aware speech reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。