arXiv:2602.02413cs.SDcs.LG2026-02

用自监督方法训练通用语音增强模型,可处理多种噪声混响。

Masked Autoencoders as Universal Speech Enhancer

  • 通过掩码自编码器学习去除多种干扰和重建频谱缺失区域。
  • 在多场景下优于基线,跨域测试也达到顶尖水平。
  • 适合语音降噪、去混响等下游任务,仅需少量标注数据微调。

监督式语音增强方法效果优异,但在实际场景中缺乏纯净语音数据。为此,本文提出一种基于掩码自编码器的通用语音增强模型,该模型不依赖特定失真类型,可同时处理多种干扰,并采用自监督方式训练。预训练阶段通过增强堆栈向含噪输入添加额外失真,模型学习移除这些新增失真并重建频谱掩码区域。预训练特征用于小样本配对数据的下游任务微调,如降噪与去混响。我们探索了不同增强策略(单/多人声)及输入特征表示(如$\log1p$压缩)对预训练特征和下游性能的影响。结果表明,该方法不仅超越基线,还在同域与跨域评估数据集上均达到当前最优表现。

原文摘要 · Abstract (English)

Supervised speech enhancement methods have been very successful. However, in practical scenarios, there is a lack of clean speech, and self-supervised learning-based (SSL) speech enhancement methods that offer comparable enhancement performance and can be applied to other speech-related downstream applications are desired. In this work, we develop a masked autoencoder based universal speech enhancer that is agnostic to the type of distortion affecting speech, can handle multiple distortions simultaneously, and is trained in a self-supervised manner. An augmentation stack adds further distortions to the noisy input data. The masked autoencoder model learns to remove the added distortions along with reconstructing the masked regions of the spectrogram during pre-training. The pre-trained embeddings are then used by fine-tuning models trained on a small amount of paired data for specific downstream tasks. We evaluate the pre-trained features for denoising and dereverberation downstream tasks. We explore different augmentations (like single or multi-speaker) in the pre-training augmentation stack and the effect of different noisy input feature representations (like $log1p$ compression) on pre-trained embeddings and downstream fine-tuning enhancement performance. We show that the proposed method not only outperforms the baseline but also achieves state-of-the-art performance for both in-domain and out-of-domain evaluation datasets.

语音增强自监督掩码自编码器通用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。