arXiv:2506.22646eess.AScs.SD2025-06中稿 · INTERSPEECH 2025被引 8

无需说话人查询,实时自适应多说话人语音识别

Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR

  • 通过说话人激活生成特异性核,动态调整识别模型
  • 在重叠语音下实现流式场景下的最优识别性能
  • 适合需要实时处理多人对话的语音系统

我们提出一种面向流式多说话人自动语音识别(ASR)的自说话人适应方法,无需显式说话人查询。与传统需目标说话人嵌入或录音注册的方法不同,本方法通过说话人级语音活动预测,动态适配每个ASR实例。关键创新在于将通过说话人监督激活生成的说话人特异性核注入选定的ASR编码器层,实现在流式场景下对目标说话人的即时适应,即使在完全重叠语音中也能有效工作。实验表明,在离线与流式场景下均达到领先性能,验证了该自适应方法在严重语音重叠条件下有效实现聚焦说话人的识别,为复杂重叠环境下的多说话人语音识别提供了稳健解决方案。

原文摘要 · Abstract (English)

We propose a self-speaker adaptation method for streaming multi-talker automatic speech recognition (ASR) that eliminates the need for explicit speaker queries. Unlike conventional approaches requiring target speaker embeddings or enrollment audio, our technique dynamically adapts individual ASR instances through speaker-wise speech activity prediction. The key innovation involves injecting speaker-specific kernels generated via speaker supervision activations into selected ASR encoder layers. This enables instantaneous speaker adaptation to target speakers while handling fully overlapped speech even in a streaming scenario. Experiments show state-of-the-art performance in both offline and streaming scenarios, demonstrating that our self-adaptive method effectively addresses severe speech overlap through streamlined speaker-focused recognition. The results validate the proposed self-speaker adaptation approach as a robust solution for multi-talker ASR under severe overlapping speech conditions.

多说话人自适应流式识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。