用语义知识蒸馏+掩码声学建模,提升语音修复的可懂性与质量。
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
- 用预训练语义模型增强语音编码器,引导重建过程
- 通过掩码语言模型预测声学标记,还原频谱细节
- 相同算力下显著降低错误率,适合语音修复任务
语音修复旨在恢复全频带语音,保证高保真度和可懂性,应对多种失真。最近提出的生成模型MaskSR虽能提供高质量输出,但可懂性仍有提升空间。本文提出改进方案MaskSR2:利用预训练自监督教师模型预测目标语音的语义表征,增强MaskSR的语音编码器;随后,基于学习到的语义特征,使用掩码语言模型预测编码低层频谱细节的声学标记。实验表明,在保持相同模型容量与推理时间的前提下,该方法显著降低词错误率(WER),在同类模型中达到竞争力表现,同时提供更优的语音质量。消融实验验证了不同语义表示的有效性。
原文摘要 · Abstract (English)
Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains high quality but, as we show, intelligibility can be substantially improved. We do so by boosting the speech encoder component of MaskSR with predictions of semantic representations of the target speech, using a pre-trained self-supervised teacher model. Then, a masked language model is conditioned on the learned semantic features to predict acoustic tokens that encode low level spectral details of the target speech. We show that, with the same MaskSR model capacity and inference time, the proposed model, MaskSR2, significantly reduces the word error rate, a typical metric for intelligibility. MaskSR2 also achieves competitive word error rate among other models, while providing superior quality. An ablation study shows the effectiveness of various semantic representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。