提升课堂环境下的语音识别准确率,适配儿童与成人声音。
Speaker Verification Under Real Classroom Conditions for English Speech

- 用自监督学习+微调的两阶段策略训练语音验证模型
- 相比基线模型,错误率降低23.99%以上
- 适合教育场景中需区分学生身份的AI应用
开发在教室噪声环境下鲁棒且适用于儿童与成人说话者的语音验证(SV)模型,对支持教育环境的AI工具至关重要。本研究采用包含部分说话人身份标注的真实英语课堂数据集,多数录音未标注。通过适配WavLM-TDNN模型于教室语音验证任务,在五折交叉验证中,相比ECAPA-TDNN基线模型实现平均23.99%的等错误率(EER)相对降低,相比在课堂数据上训练的ECAPA-TDNN模型实现6.32%的相对降低。此外,对比了两种训练策略:纯自监督学习(SSL)与先预训练再用少量标注数据微调的两阶段方法。结果表明,两阶段策略始终优于仅使用SSL的方法,平均实现13.39%的相对EER下降。
原文摘要 · Abstract (English)
Developing speaker verification (SV) models that are robust to classroom noise and effective across both children and adult speakers is critical for AI tools supporting educational environments. In this study, we use a real-world English-speaking classrooms dataset containing partial speaker identity annotations, with most recordings remaining unlabeled. We adapt the WavLM-TDNN model for classroom SV, achieving average relative reductions in Equal Error Rate (EER) of 23.99% and 6.32% compared to the ECAPA-TDNN baseline and the ECAPA-TDNN model trained on classroom data, respectively. Additionally, we investigate two training strategies for SV in classroom settings: self-supervised learning (SSL) and a two-stage approach that first pre-trains with SSL and then fine-tunes with limited annotated data. Five-fold cross-validation demonstrates that the two-stage strategy consistently outperforms the SSL-only approach, achieving an average relative EER reduction of 13.39%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。