2万小时远场多人重叠语音数据集,助力语音分离与识别研究
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
- 构建2万小时远场多人重叠语音数据集,支持单麦克风场景
- 在WHAMR!上验证了语音分离、识别与说话人辨认的联合性能
- 为会议/派对类复杂场景提供可复现基准,适合语音系统研究者
语音处理领域正日益关注会议或鸡尾酒会等多说话人、远场条件下的复杂场景。现有方法分为多通道与单通道两类,其中单通道方法因无需麦克风阵列信息而更具通用性与便利性。本文提出一个大规模远场重叠语音数据集,旨在推动语音分离、识别与说话人辨认的研究。该数据集是解码多说话人、混响环境下“谁说了什么、何时说的”的关键资源。此外,我们还引入一个包含语音分离、识别与辨认的端到端流水线系统,作为基础基准。在WHAMR!数据集上的评估验证了所提数据的广泛适用性。
原文摘要 · Abstract (English)
The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges fall into two categories: multi-channel and single-channel solutions. Single-channel approaches, notable for their generality and convenience, do not require specific information about microphone arrays. This paper presents a large-scale far-field overlapping speech dataset, crafted to advance research in speech separation, recognition, and speaker diarization. This dataset is a critical resource for decoding ``Who said What and When'' in multi-talker, reverberant environments, a daunting challenge in the field. Additionally, we introduce a pipeline system encompassing speech separation, recognition, and diarization as a foundational benchmark. Evaluations on the WHAMR! dataset validate the broad applicability of the proposed data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。