arXiv:2409.00815cs.SDcs.AI2024-09被引 9

通过分离重叠语音编码,提升多说话人语音识别效果

Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition

  • 在编码器后加分离模块,用CTC损失提取单说话人信息
  • 在LibriMix上实现单说话人编码有效分离,提升复杂场景表现
  • 适合做多说话人语音识别的模型改进,尤其关注注意力机制

序列化输出训练(SOT)因其便捷性在多说话人自动语音识别(ASR)中受到关注。然而,仅使用注意力损失进行训练存在困难。本文提出重叠编码分离(EncSep),在编码器后插入分离模块,利用连接时序分类(CTC)损失充分提取多说话人信息。进一步提出序列化语音信息引导SOT(GEncSep),将分离后的语音流拼接,为解码阶段的注意力提供单说话人信息指导。在LibriMix数据集上的实验表明,单说话人编码可从重叠编码中有效分离,且CTC损失有助于在复杂场景下优化编码表示。GEncSep进一步提升了性能。

原文摘要 · Abstract (English)

Serialized output training (SOT) attracts increasing attention due to its convenience and flexibility for multi-speaker automatic speech recognition (ASR). However, it is not easy to train with attention loss only. In this paper, we propose the overlapped encoding separation (EncSep) to fully utilize the benefits of the connectionist temporal classification (CTC) and attention hybrid loss. This additional separator is inserted after the encoder to extract the multi-speaker information with CTC losses. Furthermore, we propose the serialized speech information guidance SOT (GEncSep) to further utilize the separated encodings. The separated streams are concatenated to provide single-speaker information to guide attention during decoding. The experimental results on LibriMix show that the single-speaker encoding can be separated from the overlapped encoding. The CTC loss helps to improve the encoder representation under complex scenarios. GEncSep further improved performance.

多说话人识别语音分离CTC-Attention

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。