用吸引子机制同时分离多人多语句语音,还能自动估人数。
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
- 引入吸引子模块,动态估计说话人数量并检测活动
- 在混响噪声环境下准确分离语音,无论人数已知与否
- 适合复杂场景下的实时语音分离,如会议记录、听障辅助
本文解决单通道语音分离问题,即说话人数未知且每位说话人可能有多段发言。提出一种融合吸引子模块的分离模型,可同步完成分离、动态估计说话人数量及检测个体说话活动。该系统通过吸引子架构有效结合局部与全局时序建模,在多语句场景中表现更优。为评估在混响和噪声条件下的性能,使用Librispeech语音信号与WHAM!噪声信号合成多说话人多语句数据集。结果表明,该系统能准确估计源数量,有效检测源活动,并在已知和未知源数情况下将对应语句正确分离输出。
原文摘要 · Abstract (English)
This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously performs separation, dynamically estimates the number of speakers, and detects individual speaker activities by integrating an attractor module. The proposed system outperforms existing methods by introducing an attractor-based architecture that effectively combines local and global temporal modeling for multi-utterance scenarios. To evaluate the method in reverberant and noisy conditions, a multi-speaker multi-utterance dataset was synthesized by combining Librispeech speech signals with WHAM! noise signals. The results demonstrate that the proposed system accurately estimates the number of sources. The system effectively detects source activities and separates the corresponding utterances into correct outputs in both known and unknown source count scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。