arXiv:2510.12485eess.AS2025-10中稿 · ICASSP2026被引 1

改进的VAE框架提升单通道语音增强泛化能力

I-DCCRN-VAE: An Improved Deep Representation Learning Framework for Complex VAE-based Single-channel Speech Enhancement

  • 移除预训练VAE跳接,生成更丰富的语噪潜在表示
  • 采用β-VAE平衡重构与潜在空间正则,提升性能
  • 无需对抗训练,简化流程仍保持优异效果

近期提出的基于DCCRN架构的复数变分自编码器(VAE)单通道语音增强系统中,噪声抑制VAE(NSVAE)利用预训练的干净语音和噪声VAE及跳接结构,从含噪语音中提取清晰语音表示。本文通过三项改进:1)移除预训练VAE中的跳接以促进更具信息量的语音与噪声潜在表示;2)在预训练中使用β-VAE,更好平衡重构与潜在空间正则化;3)使NSVAE同时生成语音与噪声潜在表示。实验表明,该系统在匹配数据集DNS3上性能与DCCRN和DCCRN-VAE相当,但在不匹配数据集WSJ0-QUT与Voicebank-DEMEND上表现更优,证明其更强泛化能力。消融研究显示,采用经典微调即可达到相似性能,无需对抗训练,简化了训练流程。

原文摘要 · Abstract (English)

Recently, a complex variational autoencoder (VAE)-based single-channel speech enhancement system based on the DCCRN architecture has been proposed. In this system, a noise suppression VAE (NSVAE) learns to extract clean speech representations from noisy speech using pretrained clean speech and noise VAEs with skip connections. In this paper, we improve DCCRN-VAE by incorporating three key modifications: 1) removing the skip connections in the pretrained VAEs to encourage more informative speech and noise latent representations; 2) using $β$-VAE in pretraining to better balance reconstruction and latent space regularization; and 3) a NSVAE generating both speech and noise latent representations. Experiments show that the proposed system achieves comparable performance as the DCCRN and DCCRN-VAE baselines on the matched DNS3 dataset but outperforms the baselines on mismatched datasets (WSJ0-QUT, Voicebank-DEMEND), demonstrating improved generalization ability. In addition, an ablation study shows that a similar performance can be achieved with classical fine-tuning instead of adversarial training, resulting in a simpler training pipeline.

语音增强VAE深度学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。