用多说话人识别预训练,让语音分 speaker 更准更轻量
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
- 用混合语音识别多个说话人作为预训练任务
- 无需模拟对话数据,仍可实现高精度轻量模型
- 适合资源有限但需高效分 speaker 的场景
端到端语音分说话人通过并行估计多个说话人的语音活动,实现准确的重叠感知分说话人。该方法依赖大量标注对话数据,而真实数据难以满足需求。传统做法使用大规模模拟数据预训练,但需巨大存储与计算资源,且模拟对话真实性难保证。本文提出一种新思路:预训练模型从完全重叠的语音混合中识别多个说话人,替代传统分说话人模型的预训练。该方法无需构建大规模模拟对话数据,直接利用大规模说话人识别数据集进行训练。实验表明,该方法可实现高精度、轻量级的本地分说话人模型,且无需模拟数据。
原文摘要 · Abstract (English)
End-to-end speaker diarization enables accurate overlap-aware diarization by jointly estimating multiple speakers' speech activities in parallel. This approach is data-hungry, requiring a large amount of labeled conversational data, which cannot be fully obtained from real datasets alone. To address this issue, large-scale simulated data is often used for pretraining, but it requires enormous storage and I/O capacity, and simulating data that closely resembles real conversations remains challenging. In this paper, we propose pretraining a model to identify multiple speakers from an input fully overlapped mixture as an alternative to pretraining a diarization model. This method eliminates the need to prepare a large-scale simulated dataset while leveraging large-scale speaker recognition datasets for training. Through comprehensive experiments, we demonstrate that the proposed method enables a highly accurate yet lightweight local diarization model without simulated conversational data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。