用8通道肌电信号提升嘈杂环境下的语音清晰度
Multi-modal Speech Enhancement with Limited Electromyography Channels
- 仅用8通道肌电信号与声学信号融合增强语音
- 在极低信噪比下,语音质量提升0.527(PESQ)
- 适合噪声环境复杂、设备受限的语音应用
语音增强(SE)旨在提升各类语音应用中语音信号的清晰度、可懂性和质量。然而,空气传导语音(AC)在低信噪比和非平稳噪声环境下极易受干扰。融合多模态信息有助于改善此类挑战性场景中的语音表现。肌电图(EMG)信号能捕捉发声时的肌肉活动,具有抗噪声优势,适用于恶劣条件下的语音增强。以往基于EMG的方法通常需要35个通道,限制了实际应用。为此,我们提出一种新方法,仅使用8通道EMG信号与声学信号,结合改进的SEMamba网络及新增的跨模态模块。实验表明,该方法在语音质量和可懂性上显著优于传统方法,尤其在极端低信噪比条件下表现突出。相比仅使用声学信号的基线方法,在匹配低信噪比条件下,我们的方法在PESQ上提升了0.235;在不匹配条件下提升0.527,充分体现了其鲁棒性。
原文摘要 · Abstract (English)
Speech enhancement (SE) aims to improve the clarity, intelligibility, and quality of speech signals for various speech enabled applications. However, air-conducted (AC) speech is highly susceptible to ambient noise, particularly in low signal-to-noise ratio (SNR) and non-stationary noise environments. Incorporating multi-modal information has shown promise in enhancing speech in such challenging scenarios. Electromyography (EMG) signals, which capture muscle activity during speech production, offer noise-resistant properties beneficial for SE in adverse conditions. Most previous EMG-based SE methods required 35 EMG channels, limiting their practicality. To address this, we propose a novel method that considers only 8-channel EMG signals with acoustic signals using a modified SEMamba network with added cross-modality modules. Our experiments demonstrate substantial improvements in speech quality and intelligibility over traditional approaches, especially in extremely low SNR settings. Notably, compared to the SE (AC) approach, our method achieves a significant PESQ gain of 0.235 under matched low SNR conditions and 0.527 under mismatched conditions, highlighting its robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。