用多阶段模型提升嘈杂教室中师生语音分离的准确率
Multi-Stage Speaker Diarization for Noisy Classrooms
- 分阶段处理:先降噪再检测语音,结合ASR时间戳优化说话人识别
- 降噪后错误率降至17%(师生分离)和45%(所有说话人)
- 适合教育场景语音分析,尤其对儿童语音识别有显著提升
说话人分离旨在识别音频中'谁在何时说话',对理解课堂互动至关重要。但教室环境存在录音质量差、背景噪声大、语音重叠及儿童声音捕捉困难等挑战。本研究评估了Nvidia NeMo语音分离流水线在多阶段模型中的表现,考察降噪对分离准确率的影响,并比较多种语音活动检测(VAD)模型,包括基于自监督变换器的帧级VAD。还探索了融合自动语音识别(ASR)词级时间戳与帧级VAD预测的混合方法。在两个英语课堂数据集上进行实验,分别实现教师与学生分离及所有说话人分离。结果表明,降噪显著降低语音遗漏率,从而提升说话人分离误差率(DER)。在降噪与原始数据联合训练下,嘈杂环境中性能大幅提高。混合VAD方法进一步改善语音检测,教师-学生分离实验中达到最低17%的DER,全说话人分离为45%。但语音活动检测与说话人混淆之间存在权衡。研究证明多阶段模型与引入ASR信息对提升嘈杂课堂环境下的说话人分离有效。
原文摘要 · Abstract (English)
Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording quality, high levels of background noise, overlapping speech, and the difficulty of accurately capturing children's voices. This study investigates the effectiveness of multi-stage diarization models using Nvidia's NeMo diarization pipeline. We assess the impact of denoising on diarization accuracy and compare various voice activity detection (VAD) models, including self-supervised transformer-based frame-wise VAD models. We also explore a hybrid VAD approach that integrates Automatic Speech Recognition (ASR) word-level timestamps with frame-level VAD predictions. We conduct experiments using two datasets from English speaking classrooms to separate teacher vs. student speech and to separate all speakers. Our results show that denoising significantly improves the Diarization Error Rate (DER) by reducing the rate of missed speech. Additionally, training on both denoised and noisy datasets leads to substantial performance gains in noisy conditions. The hybrid VAD model leads to further improvements in speech detection, achieving a DER as low as 17% in teacher-student experiments and 45% in all-speaker experiments. However, we also identified trade-offs between voice activity detection and speaker confusion. Overall, our study highlights the effectiveness of multi-stage diarization models and integrating ASR-based information for enhancing speaker diarization in noisy classroom environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。