arXiv:2508.14115eess.AScs.AI2025-08

用知识蒸馏提升短时语音的说话人嵌入精度,降低追踪延迟。

Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings

  • 基于知识蒸馏训练短时语音嵌入模型,提升重识别能力。
  • 在双说话人混合场景下,嵌入效果优于传统方法,抗重叠性更强。
  • 支持固定块大小重分配,为低延迟系统提供可行性方案。

说话人嵌入是有助于身份分配的重要特征,可通过空间预测实现身份重分配。然而,常见嵌入提取器在短时上下文和重叠语音场景下表现不佳,导致需依赖长时上下文进行重分配,增加追踪错误风险。为此,本文提出一种基于知识蒸馏的短时语音嵌入提取方法,利用波束成形技术获取目标说话人的空间信息以减少重叠影响。研究了固定块大小的身份重分配策略(blockwise identity reassignment),推动低延迟说话人追踪系统的实现。实验表明,所提蒸馏模型在短时上下文嵌入提取中表现有效,且对语音重叠更具鲁棒性;但块级重分配结果也显示,同时发声处理仍需进一步优化。

原文摘要 · Abstract (English)

Speaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common speaker embedding extractors usually struggle with short temporal contexts and overlapping speech, which imposes long-term identity reassignment to exploit longer temporal contexts. However, this increases the probability of tracking system errors, which in turn impacts negatively on identity reassignment. To address this, we propose a Knowledge Distillation (KD) based training approach for short context speaker embedding extraction from two speaker mixtures. We leverage the spatial information of the speaker of interest using beamforming to reduce overlap. We study the feasibility of performing identity reassignment over blocks of fixed size, i.e., blockwise identity reassignment, to go towards a low-latency speaker embedding based tracking system. Results demonstrate that our distilled models are effective at short-context embedding extraction and more robust to overlap. Although, blockwise reassignment results indicate that further work is needed to handle simultaneous speech more effectively.

说话人追踪嵌入提取低延迟知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。