用语义预测提升语音增强的因果建模,效果显著。
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
- 用量化自监督特征+因果建模,融合语义信息增强语音
- 在VoiceBank+DEMAND上达2.88 PESQ,语义预测贡献关键
- 适合实时语音通信场景,对语音质量提升有实际价值
实时语音增强对在线语音通信至关重要。因果模型仅使用过去上下文进行未来信息预测,如音素延续,有助于提升性能。本文首次将自监督学习(SSL)特征与因果性结合用于语音增强。通过向量量化将因果SSL特征转换为语义标记,并利用特征逐维线性调制将它们与谱图特征融合,以估计增强掩码。同时,在多任务学习中,模型不仅编码特征,还预测未来语义标记。在VoiceBank + DEMAND数据集上的实验表明,该方法达到2.88 PESQ,且语义预测在其中起重要作用。
原文摘要 · Abstract (English)
Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme continuation, may help performing causal SE. The phonetic information is often represented by quantizing latent features of self-supervised learning (SSL) models. This work is the first to incorporate SSL features with causality into an SE model. The causal SSL features are encoded and combined with spectrogram features using feature-wise linear modulation to estimate a mask for enhancing the noisy input speech. Simultaneously, we quantize the causal SSL features using vector quantization to represent phonetic characteristics as semantic tokens. The model not only encodes SSL features but also predicts the future semantic tokens in multi-task learning (MTL). The experimental results using VoiceBank + DEMAND dataset show that our proposed method achieves 2.88 in PESQ, especially with semantic prediction MTL, in which we confirm that the semantic prediction played an important role in causal SE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。