用八麦复数频谱直接建模,提升混响环境下的双人语音分离效果。
Phase Aware Ear-Conditioned Learning for Multi-Channel Binaural Speaker Separation
- 输入原始复数STFT,跳过编码器直接进入解码器,提升重建精度。
- 在混响环境下达到12.37 dB SI-SDR、0.87 STOI,显著优于传统方法。
- 无需排列不变训练,适合固定方位双说话人场景的语音增强应用。
在混响环境中分离重叠语音需同时保留空间线索并保证分离效率。本文提出基于八麦克风的相位感知耳部条件分离网络(PEASE-8),直接以复数STFT为输入,绕过整个编码器路径,直接接入早期解码层以改善信号重建。模型采用基于SI-SDR的目标函数,端到端训练于直达声耳目标,联合完成双说话人的分离与去混响任务,且无需排列不变训练。在涵盖自由场、混响及噪声条件的空间化双说话人混合信号上,PEASE-8表现优异。在混响环境下(T60=0.6 s),实现12.37 dB SI-SDR、0.87 STOI和1.86 PESQ;在自由场条件下仍具竞争力。
原文摘要 · Abstract (English)
Separating competing speech in reverberant environments requires models that preserve spatial cues while maintaining separation efficiency. We present a Phase-aware Ear-conditioned speaker Separation network using eight microphones (PEASE-8) that consumes complex STFTs and directly introduces a raw-STFT input to the early decoder layer, bypassing the entire encoder pathway to improve reconstruction. The model is trained end-to-end with an SI-SDR-based objective against direct-path ear targets, jointly performing separation and dereverberation for two speakers in a fixed azimuth, eliminating the need for permutation invariant training. On spatialized two-speaker mixtures spanning anechoic, reverberant, and noisy conditions, PEASE-8 delivers strong separation and intelligibility. In reverberant environments, it achieves 12.37 dB SI-SDR, 0.87 STOI, and 1.86 PESQ at T60 = 0.6 s, while remaining competitive under anechoic conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。