用语音克隆生成失语症语音,解决数据少和隐私问题。
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
- 用商业平台克隆失语症与正常人语音,保持性别匹配。
- 95%准确识别性别,30%合成语音被误判为真实,效果逼真。
- 公开数据集,助力医疗AI诊断与个性化康复研究。
本研究探索语音克隆技术生成模仿失语症患者独特语音模式的合成语音,以应对言语语言病理学中的数据稀缺与隐私挑战。基于TORGO数据集,我们使用商业平台克隆了失语症与对照组说话人的语音,确保性别匹配。一名持证言语语言病理学家(SLP)对子集进行评估,结果表明:所有失语症案例均被正确识别,性别识别准确率达95%,但30%的合成样本被误判为真实语音,显示出高度逼真性。研究证实语音克隆能有效保留失语症特征,且合成语音已达到可被专业人士误判的程度。该方法在医疗领域具有重要意义,可缓解数据短缺、保护患者隐私,并提升人工智能驱动的诊断能力。通过构建多样化高质量语音数据集,语音克隆有助于提升模型泛化性、实现个性化治疗,推动失语症辅助技术发展。我们公开发布合成数据集,促进后续研究与合作,旨在开发更鲁棒的模型以改善患者预后。
原文摘要 · Abstract (English)
This study explores voice cloning to generate synthetic speech replicating the unique patterns of individuals with dysarthria. Using the TORGO dataset, we address data scarcity and privacy challenges in speech-language pathology. Our contributions include demonstrating that voice cloning preserves dysarthric speech characteristics, analyzing differences between real and synthetic data, and discussing implications for diagnostics, rehabilitation, and communication. We cloned voices from dysarthric and control speakers using a commercial platform, ensuring gender-matched synthetic voices. A licensed speech-language pathologist (SLP) evaluated a subset for dysarthria, speaker gender, and synthetic indicators. The SLP correctly identified dysarthria in all cases and speaker gender in 95% but misclassified 30% of synthetic samples as real, indicating high realism. Our results suggest synthetic speech effectively captures disordered characteristics and that voice cloning has advanced to produce high-quality data resembling real speech, even to trained professionals. This has critical implications for healthcare, where synthetic data can mitigate data scarcity, protect privacy, and enhance AI-driven diagnostics. By enabling the creation of diverse, high-quality speech datasets, voice cloning can improve generalizable models, personalize therapy, and advance assistive technologies for dysarthria. We publicly release our synthetic dataset to foster further research and collaboration, aiming to develop robust models that improve patient outcomes in speech-language pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。