用预训练语言模型提升语音降噪,无需文本输入即可增强听感。
Linguistic Knowledge Transfer Learning for Speech Enhancement
- 通过跨模态知识迁移,将语言模型的语义信息注入语音降噪模型。
- 在中英文数据集上均显著提升语音可懂度与降噪效果。
- 不依赖文本输入,适合真实场景下的语音处理应用。
语言知识在口语理解中至关重要,能为嘈杂环境中的语音感知提供语义和语法上下文。然而,多数语音增强(SE)方法主要依赖声学特征学习噪声与干净语音间的映射关系,对语言信息的整合有限。尽管已有文本引导的SE方法,但通常需显式语音-文本对齐或外部文本数据,限制了实际应用。此外,语言与声学表示的固有差异也带来对齐难题。本文提出跨模态知识迁移(CMKT)框架,利用预训练大语言模型(LLM)在训练阶段注入语言知识,推理时无需文本输入或调用LLM。同时引入可控时间偏移策略,增强模型鲁棒性。实验表明,CMKT在多种SE架构和LLM嵌入下均优于基线模型;在中英文数据集上的表现验证其跨语言有效性;即使无文本数据,仍具显著提升,证明其在真实场景中的实用性。通过弥合语言与声学模态的鸿沟,CMKT为语音增强提供了可扩展、创新的知识融合方案。
原文摘要 · Abstract (English)
Linguistic knowledge plays a crucial role in spoken language comprehension. It provides essential semantic and syntactic context for speech perception in noisy environments. However, most speech enhancement (SE) methods predominantly rely on acoustic features to learn the mapping relationship between noisy and clean speech, with limited exploration of linguistic integration. While text-informed SE approaches have been investigated, they often require explicit speech-text alignment or externally provided textual data, constraining their practicality in real-world scenarios. Additionally, using text as input poses challenges in aligning linguistic and acoustic representations due to their inherent differences. In this study, we propose the Cross-Modality Knowledge Transfer (CMKT) learning framework, which leverages pre-trained large language models (LLMs) to infuse linguistic knowledge into SE models without requiring text input or LLMs during inference. Furthermore, we introduce a misalignment strategy to improve knowledge transfer. This strategy applies controlled temporal shifts, encouraging the model to learn more robust representations. Experimental evaluations demonstrate that CMKT consistently outperforms baseline models across various SE architectures and LLM embeddings, highlighting its adaptability to different configurations. Additionally, results on Mandarin and English datasets confirm its effectiveness across diverse linguistic conditions, further validating its robustness. Moreover, CMKT remains effective even in scenarios without textual data, underscoring its practicality for real-world applications. By bridging the gap between linguistic and acoustic modalities, CMKT offers a scalable and innovative solution for integrating linguistic knowledge into SE models, leading to substantial improvements in both intelligibility and enhancement performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。