用语音合成的人类感知评分提升失语症语音评估效果
Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
- 引入语音合成中的人类感知标注作为外部知识
- 在自监督模型上显著提升失语症语音评估性能
- 适合语音病理评估与跨域迁移学习研究者
失语症是一种以语音可理解性下降和沟通效率降低为特征的言语障碍。自动失语症评估为帕金森病、阿尔茨海默病和中风等神经疾病诊断与治疗提供了可扩展、低成本的支持方案。本研究探索利用语音合成评估中的人类感知标注作为可靠的域外知识,用于失语症语音评估。实验结果表明,此类监督可使自监督预训练模型获得一致且显著的性能提升。研究提示,与人类判断对齐的语音合成感知评分是失语症建模的重要资源,支持有效的跨域知识迁移。
原文摘要 · Abstract (English)
Dysarthria is a speech disorder characterized by impaired intelligibility and reduced communicative effectiveness. Automatic dysarthria assessment provides a scalable, cost-effective approach for supporting the diagnosis and treatment of neurological conditions such as Parkinson's disease, Alzheimer's disease, and stroke. This study investigates leveraging human perceptual annotations from speech synthesis assessment as reliable out-of-domain knowledge for dysarthric speech assessment. Experimental results suggest that such supervision can yield consistent and substantial performance improvements in self-supervised learning pre-trained models. These findings suggest that perceptual ratings aligned with human judgments from speech synthesis evaluations represent valuable resources for dysarthric speech modeling, enabling effective cross-domain knowledge transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。