用wav2vec2.0特征引导扩散模型,提升语音增强效果
Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework

- 用FiLM在U-Net瓶颈注入wav2vec2.0特征,实现语言信息引导
- 在VoiceBank-DEMAND和LibriMix上PESQ提升0.4,其他指标也更优
- 适合做语音增强且关注语义一致性的研究者参考
扩散模型在语音增强中展现潜力,但缺乏语言指导。本文将基于wav2vec 2.0的特征作为条件,通过特征调制(FiLM)注入到U-Net瓶颈,使降噪过程受原始语音的音素表征约束。采用冻结的wav2vec 2.0编码器提取特征,由可学习的FiLM生成器产生缩放与偏移参数,对瓶颈层进行调制,计算开销极小。受线性高斯状态空间模型下最优贝叶斯因果估计启发,使用指数平滑聚合FiLM系数以实现时序压缩。在VoiceBank-DEMAND和LibriMix数据集上的评估显示,该方法在PESQ、STOI、SI-SDR和DNSMOS上均优于无条件基线,其中PESQ稳定提升0.4,表明自监督表示能有效引导基于扩散的语音增强。
原文摘要 · Abstract (English)
Diffusion models show potential for speech enhancement but lack linguistic guidance. We condition a diffusion-based model on wav2vec 2.0 features from noisy input, injected at the U-Net bottleneck via Feature-wise Linear Modulation (FiLM). Phonetic representations from wav2vec 2.0 features of degraded speech, anchor the reverse diffusion process. While a frozen wav2vec 2.0 encoder extracts features, a learned FiLM generator produces scale and shift parameters modulating the bottleneck with minimal overhead. Motivated by the optimal Bayesian causal estimator under a linear-Gaussian state-space model, FiLM coefficients are aggregated via exponential smoothing for temporal compression. Evaluation on VoiceBank-DEMAND and LibriMix shows competitive performance against the unconditioned baseline in PESQ, STOI, SI-SDR and DNSMOS. We consistently record an improvement of 0.4 on PESQ score, suggesting self-supervised representations effectively condition diffusion-based speech enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。