arXiv:2510.03630eess.AScs.SD2025-10被引 2

让多人语音识别更快:用新方法降低推理成本。

Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams

  • 将说话人特定活动信号转为无特定说话人的双流信号,解耦推理开销与说话人数量。
  • 在AMI和ICSI数据集上推理速度提升显著,性能仍保持竞争力。
  • 兼容现有系统,适合需要高效多人语音识别的场景。

多说话人自动语音识别(ASR)的常见训练范式是利用说话人活动信号,将单说话人ASR模型适配到重叠语音场景。然而,这类系统需为每个说话人运行一次ASR模型,导致推理成本随说话人数线性增长,限制了实际应用。本文提出一种新方法,通过将说话人特定的活动输出转换为两个说话人无关的活动流,实现推理成本与说话人数量解耦。核心挑战在于,直接合并活动信号会严重损害识别性能,因为预训练的ASR模型假设输入为连续、单一说话人的语音。为此,我们设计了新的启发式策略,以保留对话连贯性并兼容现有系统。实验表明,该方法可与Diarization-Conditioned Whisper(DiCoW)结合,在AMI和ICSI会议数据集上大幅降低运行时间,同时保持优异性能。

原文摘要 · Abstract (English)

An increasingly common training paradigm for multi-talker automatic speech recognition (ASR) is to use speaker activity signals to adapt single-speaker ASR models for overlapping speech. Although effective, these systems require running the ASR model once per speaker, resulting in inference costs that scale with the number of speakers and limiting their practicality. In this work, we propose a method that decouples the inference cost of activity-conditioned ASR systems from the number of speakers by converting speaker-specific activity outputs into two speaker-agnostic streams. A central challenge is that naïvely merging speaker activities into streams significantly degrades recognition, since pretrained ASR models assume contiguous, single-speaker inputs. To address this, we design new heuristics aimed at preserving conversational continuity and maintaining compatibility with existing systems. We show that our approach is compatible with Diarization-Conditioned Whisper (DiCoW) to greatly reduce runtimes on the AMI and ICSI meeting datasets while retaining competitive performance.

语音识别多说话人效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。