将说话人嵌入、语音活动检测与重叠语音检测联合训练,提升效率并简化流程。
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
- 三模块联合训练,共享模型参数,避免独立运行
- 推理速度显著提升,性能仍具竞争力
- 为端到端说话人聚类系统铺路,适合追求高效部署者
尽管当前主流是端到端说话人分离系统,但由语音活动检测(VAD)、说话人嵌入提取与聚类、以及重叠语音检测(OSD)及处理组成的模块化系统在多数情况下仍具备竞争性表现。然而,其主要缺点是各模块需独立运行和训练。本文提出一种联合训练方法,使模型同时生成说话人嵌入、执行VAD与OSD,实现接近标准方法的性能,同时推理时间大幅缩减。此外,联合推理简化了整体流程,推动构建基于聚类的统一端到端系统,可针对说话人分离目标进行优化。
原文摘要 · Abstract (English)
In spite of the popularity of end-to-end diarization systems nowadays, modular systems comprised of voice activity detection (VAD), speaker embedding extraction plus clustering, and overlapped speech detection (OSD) plus handling still attain competitive performance in many conditions. However, one of the main drawbacks of modular systems is the need to run (and train) different modules independently. In this work, we propose an approach to jointly train a model to produce speaker embeddings, VAD and OSD simultaneously and reach competitive performance at a fraction of the inference time of a standard approach. Furthermore, the joint inference leads to a simplified overall pipeline which brings us one step closer to a unified clustering-based method that can be trained end-to-end towards a diarization-specific objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。