arXiv:2409.12352eess.AScs.SD2024-09被引 7

用元信息拼接实现端到端多人语音识别,无需复杂滤波

META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR

  • 通过拼接说话人元信息动态屏蔽编码器激活
  • 在多人和目标说话人识别任务上表现媲美现有方法
  • 统一架构同时支持双任务,简化传统复杂过滤机制

我们提出一种新型端到端多人语音识别(ASR)框架,可同时实现多人(MS-ASR)和目标说话人(TS-ASR)识别。模型以全端到端方式训练,利用预训练说话人聚类模块提供的说话人监督信号。提出一种直观有效的激活掩码方法:将说话人监督模块输出用于屏蔽ASR编码器的激活,该技术称为Meta-Cat(元信息拼接),适用于两类任务。实验表明,该架构在两种任务上均达到具有竞争力的性能,无需依赖传统的神经掩码估计或音频/特征级掩码方法。此外,初步展示了可高效处理双任务的统一模型。因此,本工作证明:通过精简架构即可实现鲁棒的端到端多人语音识别,无需以往研究中复杂的说话人过滤机制。

原文摘要 · Abstract (English)

We propose a novel end-to-end multi-talker automatic speech recognition (ASR) framework that enables both multi-speaker (MS) ASR and target-speaker (TS) ASR. Our proposed model is trained in a fully end-to-end manner, incorporating speaker supervision from a pre-trained speaker diarization module. We introduce an intuitive yet effective method for masking ASR encoder activations using output from the speaker supervision module, a technique we term Meta-Cat (meta-information concatenation), that can be applied to both MS-ASR and TS-ASR. Our results demonstrate that the proposed architecture achieves competitive performance in both MS-ASR and TS-ASR tasks, without the need for traditional methods, such as neural mask estimation or masking at the audio or feature level. Furthermore, we demonstrate a glimpse of a unified dual-task model which can efficiently handle both MS-ASR and TS-ASR tasks. Thus, this work illustrates that a robust end-to-end multi-talker ASR framework can be implemented with a streamlined architecture, obviating the need for the complex speaker filtering mechanisms employed in previous studies.

语音识别多说话人端到端说话人分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。