用时域一致性提升语音去混响效果,兼顾实时性与清晰度
Dereverberation Using Binary Residual Masking with Time-Domain Consistency
- 基于STFT域的残差掩码预测,用U-Net模型抑制晚期混响
- 混合损失函数使去混响后语音保留直接成分,提升听感质量
- 适合实时语音/演唱场景,延迟低、计算效率高
语音去混响在音频处理中仍是挑战性任务,尤其对实时应用要求精度与效率并重。传统深度学习方法常因抑制混响而损害语音清晰度,而联合预测幅度和相位的方法则计算开销大。本文提出一种基于短时傅里叶变换(STFT)域残差掩码预测的实时去混响框架。采用U-Net结构训练模型,估计残差混响掩码以抑制晚期反射声,同时保留直达语音成分。通过结合二元交叉熵、残差幅度重建与时域一致性三项损失,进一步提升抑制精度与听觉感知质量。该方法实现低延迟去混响,适用于真实场景下的语音与歌唱应用。
原文摘要 · Abstract (English)
Vocal dereverberation remains a challenging task in audio processing, particularly for real-time applications where both accuracy and efficiency are crucial. Traditional deep learning approaches often struggle to suppress reverberation without degrading vocal clarity, while recent methods that jointly predict magnitude and phase have significant computational cost. We propose a real-time dereverberation framework based on residual mask prediction in the short-time Fourier transform (STFT) domain. A U-Net architecture is trained to estimate a residual reverberation mask that suppresses late reflections while preserving direct speech components. A hybrid objective combining binary cross-entropy, residual magnitude reconstruction, and time-domain consistency further encourages both accurate suppression and perceptual quality. Together, these components enable low-latency dereverberation suitable for real-world speech and singing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。