将梅尔频域引入在线多通道语音增强,显著降低计算量。
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
- 用STFT转梅尔模块压缩多通道频谱,再在梅尔域直接处理
- 计算复杂度降60%,语音增强和语音识别效果相当
- 适合对实时性要求高的语音系统部署
在线多通道语音增强近年备受关注。尽管梅尔频域更契合人耳感知且计算高效,但现有工作大多仍基于线性频域。为此,本文提出梅尔尺度框架Mel-McNet,包含两个核心组件:一个高效的STFT-to-Mel模块,将多通道STFT特征压缩为梅尔频域表示;一个改进的McNet主干网络,直接在梅尔域生成增强后的对数梅尔谱。这些谱图可直接输入声码器重建波形或供自动语音识别系统使用。在CHiME-3数据集上的实验表明,Mel-McNet在保持与原McNet相当的增强和语音识别性能的同时,计算复杂度降低了60%。其性能还优于其他主流方法,验证了梅尔尺度语音增强的潜力。
原文摘要 · Abstract (English)
Online multichannel speech enhancement has been intensively studied recently. Though Mel-scale frequency is more matched with human auditory perception and computationally efficient than linear frequency, few works are implemented in a Mel-frequency domain. To this end, this work proposes a Mel-scale framework (namely Mel-McNet). It processes spectral and spatial information with two key components: an effective STFT-to-Mel module compressing multi-channel STFT features into Mel-frequency representations, and a modified McNet backbone directly operating in the Mel domain to generate enhanced LogMel spectra. The spectra can be directly fed to vocoders for waveform reconstruction or ASR systems for transcription. Experiments on CHiME-3 show that Mel-McNet can reduce computational complexity by 60% while maintaining comparable enhancement and ASR performance to the original McNet. Mel-McNet also outperforms other SOTA methods, verifying the potential of Mel-scale speech enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。