arXiv:2606.18645eess.AScs.AI2026-06

用语音合成评分数据提升失语症语音严重度评估效果。

Augmenting Dysarthric Speech Severity Assessment with MOS Supervision

论文配图:Augmenting Dysarthric Speech Severity Assessment with MOS Supervision
图 1 · 摘自论文原文
  • 用合成语音的MOS评分数据辅助训练失语症评估模型。
  • 微调可同时提升清晰度与自然度预测性能。
  • 适合需要减少临床标注依赖的研究者使用。

失语症是一种导致语音可理解性与沟通效率降低的言语障碍。自动化的语句级失语症语音评估可支持大规模语音监测与治疗分析,但其训练受限于临床标注数据稀缺。本文提出利用语音合成评估数据(来自QualiSpeech语料库中人工标注的均值意见分数,MOS)来增强失语症评估能力。实验表明,在合成语音评估数据上进行微调能持续提升清晰度与自然度预测性能,而联合训练主要在自然度上带来增益。结果表明,合成伪影与失语症语音在感知上具有共性,语音合成评估语料库可作为实用的数据增强来源,减少对稀缺临床标注的依赖。

原文摘要 · Abstract (English)

Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness. Automatic utterance-level assessment of dysarthric speech can support scalable speech monitoring and therapy-related analysis. Yet training such systems is bottlenecked by the scarcity of clinically annotated dysarthric speech. This work proposes to augment dysarthric speech assessment using data from speech synthesis evaluations, specifically human-annotated utterances with Mean Opinion Score (MOS) labels from the QualiSpeech corpus. Experiments show that fine-tuning on speech synthesis assessment data consistently improves performance on both intelligibility and naturalness prediction, while joint training yields gains primarily on naturalness. These results suggest that synthesis artifacts and dysarthric speech share perceptual commonalities, and speech synthesis evaluation corpora offer a practical augmentation source that reduces reliance on scarce clinical annotations.

语音评估失语症数据增强MOS评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。