NTT提出多说话人远场语音识别系统,提升复杂场景下的识别准确率。
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
- 端到端语音聚类结合麦克风阵列选择,实现高精度说话人分离。
- 在真实会议场景中相对基线提升63%的宏平均词错误率性能。
- 适合研究远场语音识别、多说话人场景处理的工程师与学者。
本文介绍我们为CHiME-8挑战赛的远场多说话人自动语音识别(DASR)任务1设计的系统。该系统完成说话人计数、语音分离和语音识别,适用于从餐厅聚会至专业会议等多样录制环境,支持2至8个说话人。流程按语音分离、语音增强、语音识别顺序执行,但引入多项关键改进:首先,采用基于向量聚类的端到端说话人分离(EEND-VC),结合多通道说话人计数与目标说话人语音活动检测(TS-VAD);其次,提出新型麦克风选择策略,优化分布式麦克风间的选优机制,并改进波束成形方法;最后,利用Whisper与WavLM等语音基础模型构建多个识别模型。报告了提交挑战赛的结果及后续更新结果,最强系统相较基线实现63%的相对宏平均词错误率(macro tcpWER)提升,在几何无关系统中优于挑战赛最佳结果,尤其在NOTSOFAR-1会议数据集上表现突出。
原文摘要 · Abstract (English)
In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker counting, diarization, and ASR. It handles various recording conditions, from diner parties to professional meetings and from two to eight speakers. We perform diarization first, followed by speech enhancement, and then ASR as the challenge baseline. However, we introduced several key refinements. First, we derived a powerful speaker diarization relying on end-to-end speaker diarization with vector clustering (EEND-VC), multi-channel speaker counting using enhanced embeddings from EEND-VC, and target-speaker voice activity detection (TS-VAD). For speech enhancement, we introduced a novel microphone selection rule to better select the most relevant microphones among the distributed microphones and investigated improvements to beamforming. Finally, for ASR, we developed several models exploiting Whisper and WavLM speech foundation models. We present the results we submitted to the challenge and updated results we obtained afterward. Our strongest system achieves a 63% relative macro tcpWER improvement over the baseline and outperforms the challenge best results on the NOTSOFAR-1 meeting evaluation data among geometry-independent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。