通过置信度引导的伪标签逐步优化,提升老年语音识别准确率。
Confidence Score Guided Incremental and Speaker Adaptive Pseudo-Labeling for Semi-Supervised Elderly Speech Recognition

- 根据置信度高低分批引入未标注数据,实现渐进式学习
- 在英语和粤语老年语料上分别降低1.45%和2.27%的词错误率
- 适配不同说话人特征,尤其适合跨说话人老年语音识别任务
本文提出一种基于置信度引导的增量式、说话人自适应伪标签方法,用于半监督老年语音识别。通过设计置信度估计模块,对未转录数据进行可靠性排序,实现从高到低置信度数据的渐进式引入,形成课程学习路径。同时,利用可学习提示词进行说话人自适应训练,捕捉个体差异特征。在英语DementiaBank Pitt与粤语JCCOCC MoCA老年语音数据集上的实验表明,相比不使用置信度引导或说话人自适应的基线方法,本方法在词错误率(WER)和字符错误率(CER)上分别实现了1.45%和2.27%的绝对降低(相对降低6.21%和6.98%),效果显著。
原文摘要 · Abstract (English)
This paper proposes a novel confidence score guided incremental and speaker adaptive pseudo-labeling approach for semi-supervised elderly speech recognition. It facilitates higher-quality pseudo-label selection and progressive refinement, while also mitigating speaker heterogeneity. A confidence estimation module is designed to rank the reliability of untranscribed data, enabling a curriculum learning trajectory that progressively folds in unlabeled data subsets from high to low confidence. Speaker-specific characteristics are captured through speaker adaptive training with learnable prompts. Experiments on the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets suggest that the proposed method outperforms the semi-supervised baseline using no confidence scores guided incremental or speaker adaptive pseudo-labeling by statistically significant word error rate (WER) or character error rate (CER) reductions of 1.45% and 2.27% absolute (6.21% and 6.98% relative).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。