用语音重建技术修复儿童发音障碍,保留声音特征并提升评估效率
Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation
- 基于风格解耦的声学建模,专为有发音障碍的儿童设计
- 重建后词准确率提升,且语音身份保持率显著改善
- 可自动评估发音纠正效果,适合临床语音评估场景
我们提出ChiReSSD,一种针对儿童发音障碍(SSD)的语音重建框架,在保留说话者身份的同时抑制错误发音。与以往基于健康成人语音训练的方法不同,该方法特别关注儿童的音高和语调特征。在STAR数据集上的评估显示,词汇准确率和说话人身份保留均有显著提升。我们还自动预测原始与重建语音中的音素内容,修正辅音比例与正确辅音比例(PCC)相当,该指标是临床语音评估的重要标准。自动标注与人工标注之间的皮尔逊相关系数达0.63,表明可大幅降低手动转录负担。此外,在TORGO数据集上的实验验证了其对成年运动性构音障碍语音的泛化能力。结果表明,基于风格解耦的文本到语音重建可实现跨不同临床群体的身份保持型语音重建。
原文摘要 · Abstract (English)
We present ChiReSSD, a speech reconstruction framework that preserves children speaker's identity while suppressing mispronunciations. Unlike prior approaches trained on healthy adult speech, ChiReSSD adapts to the voices of children with speech sound disorders (SSD), with particular emphasis on pitch and prosody. We evaluate our method on the STAR dataset and report substantial improvements in lexical accuracy and speaker identity preservation. Furthermore, we automatically predict the phonetic content in the original and reconstructed pairs, where the proportion of corrected consonants is comparable to the percentage of correct consonants (PCC), a clinical speech assessment metric. Our experiments show Pearson correlation of 0.63 between automatic and human expert annotations, highlighting the potential to reduce the manual transcription burden. In addition, experiments on the TORGO dataset demonstrate effective generalization for reconstructing adult dysarthric speech. Our results indicate that disentangled, style-based TTS reconstruction can provide identity-preserving speech across diverse clinical populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。