arXiv:2505.20222eess.AS2025-05中稿 · be presented at In…被引 3

用儿童语音数据微调,显著提升课堂环境下的说话人验证鲁棒性。

FT-Boosted SV: Towards Noise Robust Speaker Verification for English Speaking Classroom Environments

  • 用增强的儿童语音数据微调x-vector和ECAPA-TDNN模型。
  • 在MPT数据集上使ECAPA-TDNN的EER降低50%(提升5%),在NCTE上x-vector提升8%。
  • 适合构建抗噪课堂语音识别与身份验证系统的研究者使用。

为教育类AI工具的发展,构建对课堂噪声(如多人嘈杂声)具有鲁棒性的说话人验证(SV)系统至关重要。本文研究了使用增强儿童语音数据进行微调,以适应x-vector和ECAPA-TDNN模型在课堂环境中的表现。结果表明,该方法有效降低了两类模型在课堂及儿童语音数据集上的等错误率(EER)。尤其在MPT数据集的课堂场景中,ECAPA-TDNN模型的EER平均降低一半(提升5%),相比基线模型;在NCTE数据集的课堂场景中,x-vector模型平均提升8%。

原文摘要 · Abstract (English)

Creating Speaker Verification (SV) systems for classroom settings that are robust to classroom noises such as babble noise is crucial for the development of AI tools that assist educational environments. In this work, we study the efficacy of finetuning with augmented children datasets to adapt the x-vector and ECAPA-TDNN to classroom environments. We demonstrate that finetuning with augmented children's datasets is powerful in that regard and reduces the Equal Error Rate (EER) of x-vector and ECAPA-TDNN models for both classroom datasets and children speech datasets. Notably, this method reduces EER of the ECAPA-TDNN model on average by half (a 5 % improvement) for classrooms in the MPT dataset compared to the ECAPA-TDNN baseline model. The x-vector model shows an 8 % average improvement for classrooms in the NCTE dataset compared to its baseline.

说话人验证课堂语音噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。