一个模型同时搞定多人语音识别、分离和说话人辨认。
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
- 用共享编码器联合学习三任务表征,提升信息利用效率。
- 在Libri2Mix和Libri3Mix上分别实现1.37%和2.29%的说话人辨认错误率。
- 适合需要处理重叠语音的语音系统研发人员参考。
本文提出一种统一多说话人编码器(UME),通过共享语音基础编码器,联合学习说话人辨认(SD)、语音分离(SS)和多说话人自动语音识别(ASR)任务的表征。利用UME多层隐藏表示进行残差加权求和(RWSE),有效融合不同语义层级的信息,促进任务间的自底向上对齐。联合训练捕捉了各任务间的内在依赖关系,在重叠语音数据上显著提升整体性能。评估表明,UME在LibriMix测试集上显著优于单任务基线模型。尤其在说话人辨认任务中,达到1.37%和2.29%的错误率,超越先前研究。
原文摘要 · Abstract (English)
This paper presents a unified multi-speaker encoder (UME), a novel architecture that jointly learns representations for speaker diarization (SD), speech separation (SS), and multi-speaker automatic speech recognition (ASR) tasks using a shared speech foundational encoder. We leverage the hidden representations from multiple layers of UME as a residual weighted-sum encoding (RWSE) to effectively use information from different semantic levels, contributing to bottom-up alignment between tasks. This joint training approach captures the inherent interdependencies among the tasks, enhancing overall performance on overlapping speech data. Our evaluations demonstrate that UME substantially improves over the single-task baselines dedicated to SD, SS, and multi-speaker ASR on LibriMix evaluation sets. Notably, for SD, UME outperforms the previous studies, achieving diarization error rates of 1.37% and 2.29% on Libri2Mix and Libri3Mix evaluation sets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。