用元信息拼接实现端到端多人语音识别,无需复杂滤波
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
- 通过拼接说话人元信息动态屏蔽编码器激活
- 在多人和目标说话人识别任务上表现媲美现有方法
- 统一架构同时支持双任务,简化传统复杂过滤机制
我们提出一种新型端到端多人语音识别(ASR)框架,可同时实现多人(MS-ASR)和目标说话人(TS-ASR)识别。模型以全端到端方式训练,利用预训练说话人聚类模块提供的说话人监督信号。提出一种直观有效的激活掩码方法:将说话人监督模块输出用于屏蔽ASR编码器的激活,该技术称为Meta-Cat(元信息拼接),适用于两类任务。实验表明,该架构在两种任务上均达到具有竞争力的性能,无需依赖传统的神经掩码估计或音频/特征级掩码方法。此外,初步展示了可高效处理双任务的统一模型。因此,本工作证明:通过精简架构即可实现鲁棒的端到端多人语音识别,无需以往研究中复杂的说话人过滤机制。
原文摘要 · Abstract (English)
We propose a novel end-to-end multi-talker automatic speech recognition (ASR) framework that enables both multi-speaker (MS) ASR and target-speaker (TS) ASR. Our proposed model is trained in a fully end-to-end manner, incorporating speaker supervision from a pre-trained speaker diarization module. We introduce an intuitive yet effective method for masking ASR encoder activations using output from the speaker supervision module, a technique we term Meta-Cat (meta-information concatenation), that can be applied to both MS-ASR and TS-ASR. Our results demonstrate that the proposed architecture achieves competitive performance in both MS-ASR and TS-ASR tasks, without the need for traditional methods, such as neural mask estimation or masking at the audio or feature level. Furthermore, we demonstrate a glimpse of a unified dual-task model which can efficiently handle both MS-ASR and TS-ASR tasks. Thus, this work illustrates that a robust end-to-end multi-talker ASR framework can be implemented with a streamlined architecture, obviating the need for the complex speaker filtering mechanisms employed in previous studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。