用倒谱平滑优化语音分离中的时频掩码,减少音乐噪声。
Cepstral Smoothing of Binary Masks for Convolutive Blind Separation of Speech Mixtures
- 先用盲源分离估计二值时频掩码,再做倒谱平滑处理
- 在模拟和真实录音上均有效降低音乐噪声,提升分离效果
- 适合需要高质量语音分离的场景,如会议记录或听障辅助
本文提出一种新系统,从两个麦克风录制信号中分离出两路语音。该系统结合盲源分离技术与二值时频掩码的倒谱平滑。掩码生成分两步:首先由盲源分离算法输出信号估计二值掩码;其次对这些频谱掩码进行倒谱平滑,以减少时频掩码常带来的音乐噪声。实验在模拟房间模型生成的人工混响语音和两组真实录音上进行,评估结果表明该系统有效,具有良好的分离性能。
原文摘要 · Abstract (English)
In this paper, we propose a novel separation system for extracting two speech signals from two microphone recordings. Our system combines the blind source separation technique with cepstral smoothing of binary time-frequency masks. The last is composed of two steps. First, the two binary masks are estimated from the separated output signals of BSS algorithm. In the second step, a cepstral smoothing is applied of these spectral masks in order to reduce musical noise typically produced by time-frequency masking. Experiments were carried out with both artificially mixed speech signals using simulated room model and two real recordings. The evaluation results are promising and have shown the effectiveness of our system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。