融合多个会议识别系统,提升语音识别准确率
MOVER: Combining Multiple Meeting Recognition Systems
- 通过五阶段流程融合不同说话人分割和识别结果
- 在两个任务上分别提升9.55%和8.51%的通话错误率
- 适合需要多系统集成的会议语音识别场景
本文提出一种新的会议识别系统融合方法MOVER,用于结合不同说话人分离和语音识别输出的会议识别系统。尽管已有方法可融合说话人分离(如DOVER)或语音识别(如ROVER)系统的输出,但MOVER是首个能融合在说话人分离和语音识别两方面均存在差异的会议识别系统输出的方法。MOVER通过五阶段流程实现,包括说话人对齐、段落分组、词与时间信息合并等。在CHiME-8 DASR任务和NOTSOFAR-1多通道任务上的实验表明,该方法能有效融合多个具有差异性输出的会议识别系统,在两项任务上相对最优系统分别实现9.55%和8.51%的通话错误率(tcpWER)改进。
原文摘要 · Abstract (English)
In this paper, we propose Meeting recognizer Output Voting Error Reduction (MOVER), a novel system combination method for meeting recognition tasks. Although there are methods to combine the output of diarization (e.g., DOVER) or automatic speech recognition (ASR) systems (e.g., ROVER), MOVER is the first approach that can combine the outputs of meeting recognition systems that differ in terms of both diarization and ASR. MOVER combines hypotheses with different time intervals and speaker labels through a five-stage process that includes speaker alignment, segment grouping, word and timing combination, etc. Experimental results on the CHiME-8 DASR task and the multi-channel track of the NOTSOFAR-1 task demonstrate that MOVER can successfully combine multiple meeting recognition systems with diverse diarization and recognition outputs, achieving relative tcpWER improvements of 9.55 % and 8.51 % over the state-of-the-art systems for both tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。