arXiv:2409.05554eess.AS2024-09被引 13

针对远场语音识别,提出多说话人分离与识别系统,性能较基线提升57%。

NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge

  • 先语音分离后识别,用端到端聚类+目标说话人检测优化
  • 自研多通道说话人计数,适配不同人数场景,开发集宏平均词错误率21.3%
  • 结合强预训练模型,适合复杂环境下的多人远场语音识别任务

我们为CHiME-8挑战赛的远场自动语音识别(DASR)任务提出了一套系统。该系统采用先说话人分离的流程:首先使用端到端说话人聚类(EEND-VC)进行说话人分离,并通过目标说话人语音活动检测(TS-VAD)进行精细化处理;为应对不同数量的说话人,我们提出了新的多通道说话人计数方法;随后对基线系统进行了改进的引导源分离(GSS);最后利用多个基于强预训练模型构建的ASR系统进行融合。所提系统在开发集上实现了21.3%的宏平均词错误率(macro tcpWER),相比基线系统相对提升57%。

原文摘要 · Abstract (English)

We present a distant automatic speech recognition (DASR) system developed for the CHiME-8 DASR track. It consists of a diarization first pipeline. For diarization, we use end-to-end diarization with vector clustering (EEND-VC) followed by target speaker voice activity detection (TS-VAD) refinement. To deal with various numbers of speakers, we developed a new multi-channel speaker counting approach. We then apply guided source separation (GSS) with several improvements to the baseline system. Finally, we perform ASR using a combination of systems built from strong pre-trained models. Our proposed system achieves a macro tcpWER of 21.3 % on the dev set, which is a 57 % relative improvement over the baseline.

语音识别远场语音说话人分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。