arXiv:2506.20288eess.AScs.SD2025-06

轻量级方法实现重叠语音中目标说话人精准转录

Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR

  • 用不变模型+条件模型切换,只在重叠时激活目标说话人识别
  • 重叠段错误率从68.0%降至35.78%,计算开销仅增44%
  • 适合实时广播、会议等多说话人场景的语音服务部署

重叠语音仍是实际应用中自动语音识别(ASR)的主要挑战,尤其在动态多说话人互动的广播媒体中。本文提出一种轻量级、基于目标说话人的流式ASR扩展方法,实现在极低计算开销下对重叠语音进行实用转录。该方法结合非说话人依赖(SI)模型标准运行与说话人条件(SC)模型在重叠场景下选择性使用。重叠检测通过一个紧凑二分类器实现,基于冻结的SI模型输出,准确分割重叠段且成本极低。SC模型采用特征逐维线性调制(FiLM)融合说话人嵌入,并在合成混音数据上训练以只转录目标说话人。系统支持动态说话人追踪,复用现有模块且修改极少。在包含16%重叠的捷克电视辩论数据集上评估,重叠段词错误率(WER)由基线68.0%降至35.78%,总计算负载仅增加44%。该系统为连续语音识别服务中的重叠转录提供了高效可扩展的解决方案。

原文摘要 · Abstract (English)

Overlapping speech remains a major challenge for automatic speech recognition (ASR) in real-world applications, particularly in broadcast media with dynamic, multi-speaker interactions. We propose a light-weight, target-speaker-based extension to an existing streaming ASR system to enable practical transcription of overlapping speech with minimal computational overhead. Our approach combines a speaker-independent (SI) model for standard operation with a speaker-conditioned (SC) model selectively applied in overlapping scenarios. Overlap detection is achieved using a compact binary classifier trained on frozen SI model output, offering accurate segmentation at negligible cost. The SC model employs Feature-wise Linear Modulation (FiLM) to incorporate speaker embeddings and is trained on synthetically mixed data to transcribe only the target speaker. Our method supports dynamic speaker tracking and reuses existing modules with minimal modifications. Evaluated on a challenging set of Czech television debates with 16% overlap, the system reduced WER on overlapping segments from 68.0% (baseline) to 35.78% while increasing total computational load by only 44%. The proposed system offers an effective and scalable solution for overlap transcription in continuous ASR services.

语音识别重叠语音流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。