arXiv:2506.18714cs.SDcs.AI2025-06被引 4

改进语音增强的损失函数,让模型更关注关键频段的发音清晰度。

Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement

  • 在时频域设计可调节权重的损失函数,突出语音强或噪声大的区域
  • 实验显示语音保真度提升,尤其在辅音重建上明显改善
  • 适合需要高语音可懂度的场景,如降噪耳机、会议系统

深度学习显著提升了多通道语音增强性能,但传统损失函数(如尺度不变信干比,SDR)难以保留对音素可懂性至关重要的精细谱线索。本文提出感知启发的SDR损失变体,基于时频域并引入频率依赖加权机制,强调语音显著或噪声强烈的时频区域。研究了固定与自适应策略,包括ANSI频带重要性权重、谱幅加权及基于语音与噪声相对占比的动态加权。使用这些损失训练FaSNet多通道语音增强模型。实验表明,尽管标准指标如SDR仅略有提升,其感知加权版本却有显著改进;谱分析与音素级分析显示辅音重建更佳,表明关键声学线索得到更好保留。

原文摘要 · Abstract (English)

Recent advances in deep learning have significantly improved multichannel speech enhancement algorithms, yet conventional training loss functions such as the scale-invariant signal-to-distortion ratio (SDR) may fail to preserve fine-grained spectral cues essential for phoneme intelligibility. In this work, we propose perceptually-informed variants of the SDR loss, formulated in the time-frequency domain and modulated by frequency-dependent weighting schemes. These weights are designed to emphasize time-frequency regions where speech is prominent or where the interfering noise is particularly strong. We investigate both fixed and adaptive strategies, including ANSI band-importance weights, spectral magnitude-based weighting, and dynamic weighting based on the relative amount of speech and noise. We train the FaSNet multichannel speech enhancement model using these various losses. Experimental results show that while standard metrics such as the SDR are only marginally improved, their perceptual frequency-weighted counterparts exhibit a more substantial improvement. Besides, spectral and phoneme-level analysis indicates better consonant reconstruction, which points to a better preservation of certain acoustic cues.

语音增强深度学习音素级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。