融合神经波束成形与语音识别,提升远场会议转录准确率
Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription
- 用真实会议数据预训练神经波束成形器,提升噪声抑制能力
- 联合优化波束成形与识别模型,相对词错误率降低9%
- 适合做远场多麦克风语音识别的科研与工程人员
远场麦克风会议转录是一项挑战性任务。当前主流的端到端说话人归属语音识别(SA-ASR)架构缺乏多通道降噪与混响抑制的前端,限制了其性能。本文提出一种联合波束成形与SA-ASR的方法用于真实会议转录。首先,设计一种数据对齐与增强方法,基于真实会议数据预训练神经波束成形器。接着,对比固定、混合和全神经波束成形器作为SA-ASR前端的效果。最后,联合优化全神经波束成形器与SA-ASR模型。在真实AMI语料库上的实验表明,尽管基于多帧跨通道注意力的通道融合未能提升语音识别性能,但在固定波束成形器输出上微调SA-ASR,以及联合微调神经波束成形器与SA-ASR,分别使词错误率相对降低8%和9%。
原文摘要 · Abstract (English)
Distant-microphone meeting transcription is a challenging task. State-of-the-art end-to-end speaker-attributed automatic speech recognition (SA-ASR) architectures lack a multichannel noise and reverberation reduction front-end, which limits their performance. In this paper, we introduce a joint beamforming and SA-ASR approach for real meeting transcription. We first describe a data alignment and augmentation method to pretrain a neural beamformer on real meeting data. We then compare fixed, hybrid, and fully neural beamformers as front-ends to the SA-ASR model. Finally, we jointly optimize the fully neural beamformer and the SA-ASR model. Experiments on the real AMI corpus show that, while state-of-the-art multi-frame cross-channel attention based channel fusion fails to improve ASR performance, fine-tuning SA-ASR on the fixed beamformer's output and jointly fine-tuning SA-ASR with the neural beamformer reduce the word error rate by 8% and 9% relative, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。