arXiv:2409.04859cs.SDeess.AS2024-09被引 4

用生成模型提升语音活动检测,让系统能采样多种结果并更好推理。

Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching

  • 将二值标签映射到密集隐空间,再用流匹配生成新结果
  • 仅需两次推理就达到优秀性能,且可多次采样提升准确率
  • 首次将生成方法用于语音分说话人,适合追求多样性的研究者

语音分说话人通常被视为判别任务,采用判别方法输出固定结果。本文首次探索使用基于神经网络的生成方法进行语音分说话人。在序列到序列的目标说话人语音活动检测(Seq2Seq-TSVAD)系统中引入流匹配(Flow-Matching, FM)生成算法。实验表明,直接在原始二值标签序列空间应用生成方法无效。为此,我们提出先将二值标签序列映射到稠密隐空间,再施加生成算法,所提的Flow-TSVAD方法显著优于原Seq2Seq-TSVAD系统。此外,观察到FM算法在推理阶段收敛极快,仅需两次推理即达良好效果。作为生成模型,Flow-TSVAD可通过多次运行采样不同分说话人结果,且多实例集成进一步提升性能。

原文摘要 · Abstract (English)

Speaker diarization is typically considered a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore the use of neural network-based generative methods for speaker diarization for the first time. We implement a Flow-Matching (FM) based generative algorithm within the sequence-to-sequence target speaker voice activity detection (Seq2Seq-TSVAD) diarization system. Our experiments reveal that applying the generative method directly to the original binary label sequence space of the TS-VAD output is ineffective. To address this issue, we propose mapping the binary label sequence into a dense latent space before applying the generative algorithm and our proposed Flow-TSVAD method outperforms the Seq2Seq-TSVAD system. Additionally, we observe that the FM algorithm converges rapidly during the inference stage, requiring only two inference steps to achieve promising results. As a generative model, Flow-TSVAD allows for sampling different diarization results by running the model multiple times. Moreover, ensembling results from various sampling instances further enhances diarization performance.

语音分说话人生成模型流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。