用合成数据提升多人语音识别与说话人分离,发现不同策略效果差异大。
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
- 设计高效模拟器FastMSS,测试多种语音混合策略
- 增加语音重叠提升识别率,但会降低说话人分离精度
- 来源多样性比精确匹配领域更重要,合成数据可媲美真实数据
当前多人语音识别(MT-ASR)和说话人分离(SD)系统依赖合成数据缓解真实对话录音稀缺问题,但具体仿真方案的影响尚不明确。为弥合模拟混合与真实交互间的差距,我们研究了主流MT-ASR(DiCoW)和SD(Sortformer)系统的合成数据生成方法。通过引入高效开源模拟器FastMSS,分析了对话轮换机制、源域、声学增强及数据混合策略。结果表明:最优仿真方案高度依赖任务——增加语音重叠有利于ASR但损害SD;广泛来源多样性始终优于精确领域匹配。最终,纯合成数据训练可达到真实数据基线水平,而将模拟数据与真实录音结合,在两项任务上均显著优于仅使用真实数据的训练方式。
原文摘要 · Abstract (English)
Recent breakthroughs in multi-talker ASR (MT-ASR) and speaker diarization (SD) rely on synthetic data to mitigate the scarcity of large-scale conversational recordings, yet the impact of specific simulation choices remains poorly understood. To mind the gap between simulated mixtures and real-world interactions, we present a study of synthetic data generation for leading MT-ASR (DiCoW) and SD (Sortformer) systems. By introducing FastMSS, a highly efficient open-source simulator, we analyze turn-taking dynamics, source domain, acoustic augmentation, and data mixing strategies. Our findings reveal that optimal simulation recipes are highly task-dependent: increasing speech overlap benefits ASR but degrades diarization. Furthermore, broad source diversity consistently outperforms exact domain matching. Ultimately, synthetic-only training approaches real-data baselines, and combining simulated data with real recordings yields substantial gains over real-only training across both tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。