用语言信息提升语音增强效果,让模型更懂说话内容。
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
- 用扩散模型融合音视频与语言信息,实现跨模态知识迁移。
- 相比现有方法,语音质量显著提升,发音混淆问题减少。
- 训练后无需语言模型,适合实际部署,尤其适合多模态场景。
语音增强(SE)旨在改善嘈杂环境中的语音质量和可懂度。近期研究发现,引入视觉线索可提升音频处理性能。由于人类交流天然包含音频、视觉和语言三种模态,整合语言信息有望带来进一步改进。然而,有效弥合这些模态间差距,尤其是知识迁移过程中的挑战仍存在。本文提出一种新型多模态学习框架DLAV-SE,采用基于扩散的模型融合音频、视觉和语言信息,用于音视频语音增强(AVSE)。其中,语言模态通过预训练语言模型(PLM)建模,并在训练中通过跨模态知识迁移(CMKT)机制将语言知识传递至音视频域。训练完成后,推理阶段不再需要PLM,其知识已嵌入到AVSE模型中。大量实验表明,所提方法显著提升语音质量,减少生成伪影(如发音混淆),优于当前最优(SOTA)方法。可视化分析进一步证实,CMKT提升了输出生成质量。结果表明,扩散模型在推进AVSE方面具有潜力,而融入语言信息能进一步提升系统性能。
原文摘要 · Abstract (English)
Speech enhancement (SE) aims to improve the quality and intelligibility of speech in noisy environments. Recent studies have shown that incorporating visual cues in audio signal processing can enhance SE performance. Given that human speech communication naturally involves audio, visual, and linguistic modalities, it is reasonable to expect additional improvements by integrating linguistic information. However, effectively bridging these modality gaps, particularly during knowledge transfer remains a significant challenge. In this paper, we propose a novel multi-modal learning framework, termed DLAV-SE, which leverages a diffusion-based model integrating audio, visual, and linguistic information for audio-visual speech enhancement (AVSE). Within this framework, the linguistic modality is modeled using a pretrained language model (PLM), which transfers linguistic knowledge to the audio-visual domain through a cross-modal knowledge transfer (CMKT) mechanism during training. After training, the PLM is no longer required at inference, as its knowledge is embedded into the AVSE model through the CMKT process. We conduct a series of SE experiments to evaluate the effectiveness of our approach. Results show that the proposed DLAV-SE system significantly improves speech quality and reduces generative artifacts, such as phonetic confusion, compared to state-of-the-art (SOTA) methods. Furthermore, visualization analyses confirm that the CMKT method enhances the generation quality of the AVSE outputs. These findings highlight both the promise of diffusion-based methods for advancing AVSE and the value of incorporating linguistic information to further improve system performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。