arXiv:2502.20040eess.AScs.AI2025-02被引 9

CleanMel通过增强梅尔谱图提升语音质量与识别准确率

CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR

  • 在梅尔频率域交替使用跨频带与窄带处理,学习全频段与局部信号特征
  • 在5个英文和1个中文数据集上同时提升语音质量和语音识别性能
  • 可直接用于语音合成或语音识别,适合多场景语音增强应用

本文提出CleanMel,一种单通道梅尔谱图去噪与去混响网络,旨在同时提升语音质量与自动语音识别(ASR)性能。该网络输入为噪声与混响的麦克风录音,输出对应干净的梅尔谱图。增强后的梅尔谱图可通过神经声码器转换为语音波形,或直接用于ASR。网络采用梅尔频率域中交错的跨频带与窄带处理结构,分别学习全频段谱图模式与窄带信号特性。相较于线性频率域或时域语音增强,梅尔谱图能更紧凑地表示语音信息,利于模型学习,从而同时改善语音质量与ASR表现。在5个英文和1个中文数据集上的实验表明,所提模型显著提升了语音质量和识别性能。代码与音频示例已公开。

原文摘要 · Abstract (English)

In this work, we propose CleanMel, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) performance. The proposed network takes as input the noisy and reverberant microphone recording and predicts the corresponding clean Mel-spectrogram. The enhanced Mel-spectrogram can be either transformed to the speech waveform with a neural vocoder or directly used for ASR. The proposed network is composed of interleaved cross-band and narrow-band processing in the Mel-frequency domain, for learning the full-band spectral pattern and the narrow-band properties of signals, respectively. Compared to linear-frequency domain or time-domain speech enhancement, the key advantage of Mel-spectrogram enhancement is that Mel-frequency presents speech in a more compact way and thus is easier to learn, which will benefit both speech quality and ASR. Experimental results on five English and one Chinese datasets demonstrate a significant improvement in both speech quality and ASR performance achieved by the proposed model.Code and audio examples of our model are available online.

语音增强梅尔谱图ASR优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。