arXiv:2412.05589eess.AScs.SD2024-12中稿 · IEEE/ACM TASLP被引 12

让语音大模型只听指定说话人,解决多人重叠对话识别难题

SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR

  • 用可训练的说话人查询向量,从混杂语音中提取目标说话人特征
  • 在Libri2Mix和WSJ0-2Mix上相对降低15%和10%错误率,刷新最优纪录
  • 适合需要精准识别特定说话人的会议转录、客服记录等场景

得益于海量多源数据,语音基础模型具备强大的泛化与知识迁移能力,适用于多种下游任务。然而,其仅处理单说话人输入的局限性使其难以识别真实场景中常见的多人重叠语音。本研究探索将语音基础模型适配至目标说话人自动语音识别(TS-ASR),以消除干扰说话人影响。我们以Whisper模型为基础,对比其与现有目标说话人适配技术的融合效果,并提出新型模型SQ-Whisper:通过一组可训练查询向量,基于目标说话人注册信息捕获重叠语音中的说话人提示,引导模型提取说话人特异性特征并准确识别目标说话人转录内容。实验表明,该方法显著提升预训练语音模型在TS-ASR上的性能:相比鲁棒的TS-HuBERT模型,在Libri2Mix和WSJ0-2Mix数据集上分别实现最高15%和10%的词错误率(WER)相对下降;经数据增强后,取得14.6%(Libri2Mix测试集)和4.4%(WSJ0-2Mix测试集)的新最佳WER。此外,在真实会议数据集AMI上也持续优于其他适配方法。

原文摘要 · Abstract (English)

Benefiting from massive and diverse data sources, speech foundation models exhibit strong generalization and knowledge transfer capabilities to a wide range of downstream tasks. However, a limitation arises from their exclusive handling of single-speaker speech input, making them ineffective in recognizing multi-speaker overlapped speech, a common occurrence in real-world scenarios. In this study, we delve into the adaptation of speech foundation models to eliminate interfering speakers from overlapping speech and perform target-speaker automatic speech recognition (TS-ASR). Initially, we utilize the Whisper model as the foundation for adaptation and conduct a thorough comparison of its integration with existing target-speaker adaptation techniques. We then propose an innovative model termed Speaker-Querying Whisper (SQ-Whisper), which employs a set number of trainable queries to capture speaker prompts from overlapping speech based on target-speaker enrollment. These prompts serve to steer the model in extracting speaker-specific features and accurately recognizing target-speaker transcriptions. Experimental results demonstrate that our approach effectively adapts the pre-trained speech foundation model to TS-ASR. Compared with the robust TS-HuBERT model, the proposed SQ-Whisper significantly improves performance, yielding up to 15% and 10% relative reductions in word error rates (WERs) on the Libri2Mix and WSJ0-2Mix datasets, respectively. With data augmentation, we establish new state-of-the-art WERs of 14.6% on the Libri2Mix Test set and 4.4% on the WSJ0-2Mix Test set. Furthermore, we evaluate our model on the real-world AMI meeting dataset, which shows consistent improvement over other adaptation methods.

语音识别目标说话人混合语音Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。