arXiv:2606.21635cs.SDcs.CL2026-06中稿 · Interspeech 2026

针对语音增强中频谱细节丢失问题,提出时频加权损失函数提升辅音可懂度。

Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement

  • 基于语音存在度、信干比和谱流设计可微分时频加权机制
  • 在低信干比下显著提升中频结构重建效果,辅音识别率提高
  • 适合关注语音清晰度与发音准确性的语音增强研究者

基于信号失真比(SDR)的传统语音增强训练损失对所有时频(TF)区域一视同仁,忽略了影响特定音素可懂度的精细频谱线索。本文提出一种时频加权框架,根据局部语音存在度、语音干扰比(SIR)和谱流调节SDR目标。通过将这些因素融入可微分目标函数,该框架突出强调语音-噪声竞争激烈的时频单元,同时考虑辅音爆发等瞬态线索。实验表明,该方法提升了客观频率加权增强指标,尤其改善了辅音的音素识别准确率。频谱分析显示,在较差SIR条件下,中频结构的重建更优。

原文摘要 · Abstract (English)

Conventional training losses for speech enhancement based on the signal-to-distortion ratio (SDR) treat all time-frequency (TF) regions uniformly, overlooking the fine-grained spectral cues that are relevant to specific phoneme intelligibility. We propose a TF weighting framework that modulates the SDR objective based on local speech presence, speech-to-interference ratio (SIR), and spectral flux. By integrating these factors into a differentiable objective, the framework emphasizes TF bins with high speech-noise competition while also accounting for transient cues such as consonant bursts. Experimental results show that our approach improves objective frequency-weighted enhancement metrics, as well as phoneme recognition accuracy, particularly for consonants. Spectral analysis shows better reconstruction of mid-frequency structures at less adverse SIR.

语音增强时频加权音素识别辅音重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。