arXiv:2409.00819cs.SDcs.CL2024-09被引 8

2万小时远场多人重叠语音数据集,助力语音分离与识别研究

LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization

  • 构建2万小时远场多人重叠语音数据集,支持单麦克风场景
  • 在WHAMR!上验证了语音分离、识别与说话人辨认的联合性能
  • 为会议/派对类复杂场景提供可复现基准,适合语音系统研究者

语音处理领域正日益关注会议或鸡尾酒会等多说话人、远场条件下的复杂场景。现有方法分为多通道与单通道两类,其中单通道方法因无需麦克风阵列信息而更具通用性与便利性。本文提出一个大规模远场重叠语音数据集,旨在推动语音分离、识别与说话人辨认的研究。该数据集是解码多说话人、混响环境下“谁说了什么、何时说的”的关键资源。此外,我们还引入一个包含语音分离、识别与辨认的端到端流水线系统,作为基础基准。在WHAMR!数据集上的评估验证了所提数据的广泛适用性。

原文摘要 · Abstract (English)

The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges fall into two categories: multi-channel and single-channel solutions. Single-channel approaches, notable for their generality and convenience, do not require specific information about microphone arrays. This paper presents a large-scale far-field overlapping speech dataset, crafted to advance research in speech separation, recognition, and speaker diarization. This dataset is a critical resource for decoding ``Who said What and When'' in multi-talker, reverberant environments, a daunting challenge in the field. Additionally, we introduce a pipeline system encompassing speech separation, recognition, and diarization as a foundational benchmark. Evaluations on the WHAMR! dataset validate the broad applicability of the proposed data.

语音分离远场语音说话人辨认数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。