用伪标签提升远场语音增强,显著改善会议场景下的语音识别效果。
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
- 设计G-SpatialNet模型,优化语音分离信号质量
- 提出TLS框架生成伪标签,解决真实录音数据训练难题
- 结合多模态信息与微调策略,提升预训练语音识别性能
本文介绍了参加MISP-Meeting Challenge Track 2的系统。挑战主要源于数据集包含强背景噪声、混响、重叠语音及多样会议话题。为应对这些困难,我们(a)设计了G-SpatialNet语音增强(SE)模型,用于提升引导源分离(GSS)信号;(b)提出了TLS框架,包括时间对齐、电平对齐和信噪比滤波,用于生成真实远场音频数据的信号级伪标签,从而促进SE模型训练;(c)探索了微调策略、数据增强和多模态信息融合,以提升预训练自动语音识别(ASR)模型在会议场景中的表现。最终系统在开发集和评测集上分别取得5.44%和9.52%的字符错误率(CER),相比基线相对提升64.8%和52.6%,获得第二名。
原文摘要 · Abstract (English)
This paper presents our system for the MISP-Meeting Challenge Track 2. The primary difficulty lies in the dataset, which contains strong background noise, reverberation, overlapping speech, and diverse meeting topics. To address these issues, we (a) designed G-SpatialNet, a speech enhancement (SE) model to improve Guided Source Separation (GSS) signals; (b) proposed TLS, a framework comprising time alignment, level alignment, and signal-to-noise ratio filtering, to generate signal-level pseudo labels for real-recorded far-field audio data, thereby facilitating SE models' training; and (c) explored fine-tuning strategies, data augmentation, and multimodal information to enhance the performance of pre-trained Automatic Speech Recognition (ASR) models in meeting scenarios. Finally, our system achieved character error rates (CERs) of 5.44% and 9.52% on the Dev and Eval sets, respectively, with relative improvements of 64.8% and 52.6% over the baseline, securing second place.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。