用无标签语音数据提升构音障碍语音评估的鲁棒性
Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech
- 三阶段框架:伪标签生成+对比学习预训练+微调
- 跨病种跨语言测试平均SRCC达0.761,优于SOTA方法
- 适合医疗语音评估、无障碍技术研究者参考
构音障碍语音质量评估(DSQA)对临床诊断和包容性语音技术至关重要。然而,主观评价成本高且难以扩展,标注数据稀缺限制了稳健的客观建模。为此,我们提出一种三阶段框架,利用未标注的构音障碍语音和大规模正常语音数据集扩充训练数据。首先通过教师模型为无标签样本生成伪标签,接着采用感知标签的对比学习策略进行弱监督预训练,使模型适应多样说话人和声学环境。最后在下游DSQA任务上微调。在五个涵盖多种病因和语言的未见数据集上的实验表明,该方法具有强鲁棒性。基于Whisper的基线模型显著优于当前最优的DSQA预测器SpICE,完整框架在未见测试集上平均SRCC达到0.761。
原文摘要 · Abstract (English)
Dysarthric speech quality assessment (DSQA) is critical for clinical diagnostics and inclusive speech technologies. However, subjective evaluation is costly and difficult to scale, and the scarcity of labeled data limits robust objective modeling. To address this, we propose a three-stage framework that leverages unlabeled dysarthric speech and large-scale typical speech datasets to scale training. A teacher model first generates pseudo-labels for unlabeled samples, followed by weakly supervised pretraining using a label-aware contrastive learning strategy that exposes the model to diverse speakers and acoustic conditions. The pretrained model is then fine-tuned for the downstream DSQA task. Experiments on five unseen datasets spanning multiple etiologies and languages demonstrate the robustness of our approach. Our Whisper-based baseline significantly outperforms SOTA DSQA predictors such as SpICE, and the full framework achieves an average SRCC of 0.761 across unseen test datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。