arXiv:2409.02041eess.AScs.SD2024-09被引 16

针对复杂会议场景的语音识别挑战,优化了音源分离与识别模型。

The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge

  • 联合训练语音分离与说话人辨识,提升音频质量。
  • 多通道与单通道分别实现14.265%和22.989%的最低词错误率。
  • 适合关注实际会议环境语音处理的研究者。

本文介绍了我们在CHiME-8 NOTSOFAR-1挑战中的系统方案。该挑战数据来自多个会议室,存在高重叠率、背景噪声、说话人数不固定及自然对话风格等现实复杂性。为应对这些问题,我们在前端语音信号处理方面,采用数据驱动的联合训练方法(JDS)进行说话人分离与辨识,并结合传统引导源分离(GSS)为多通道任务提供互补信息。在后端语音识别方面,通过引入WavLM、ConvNeXt与Transformer改进Whisper模型,采用多任务训练与噪声KLD增强策略,显著提升了语音识别的鲁棒性与准确率。最终系统在多通道和单通道任务上的时间约束最小排列词错误率(tcpWER)分别为14.265%和22.989%。

原文摘要 · Abstract (English)

This technical report outlines our submission system for the CHiME-8 NOTSOFAR-1 Challenge. The primary difficulty of this challenge is the dataset recorded across various conference rooms, which captures real-world complexities such as high overlap rates, background noises, a variable number of speakers, and natural conversation styles. To address these issues, we optimized the system in several aspects: For front-end speech signal processing, we introduced a data-driven joint training method for diarization and separation (JDS) to enhance audio quality. Additionally, we also integrated traditional guided source separation (GSS) for multi-channel track to provide complementary information for the JDS. For back-end speech recognition, we enhanced Whisper with WavLM, ConvNeXt, and Transformer innovations, applying multi-task training and Noise KLD augmentation, to significantly advance ASR robustness and accuracy. Our system attained a Time-Constrained minimum Permutation Word Error Rate (tcpWER) of 14.265% and 22.989% on the CHiME-8 NOTSOFAR-1 Dev-set-2 multi-channel and single-channel tracks, respectively.

语音识别多说话人会议场景模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。