用嘴型动作增强语音降噪,让听不清的对话变清晰。
Visual-Informed Speech Enhancement Using Attention-Based Beamforming
- 结合麦克风阵列与深度网络,用嘴型识别说话人位置。
- 在低信噪比和动态说话场景下,性能优于传统方法。
- 适合嘈杂环境中需精准定位说话人的应用。
近期研究显示,引入辅助信息如语音特征或视觉线索可显著提升语音增强(SE)效果。然而,在低信噪比、高混响或涉及动态说话人、重叠语音及非平稳噪声等复杂场景中,单通道方法表现仍不理想。为此,我们提出一种新型视听神经波束成形网络(VI-NBFNet),融合麦克风阵列信号处理与深度神经网络(DNN),采用多模态输入特征。该网络利用预训练的视觉语音识别模型提取唇部运动作为输入特征,用于语音活动检测(VAD)与目标说话人识别。系统通过引入监督式端到端波束成形框架及注意力机制,能够有效处理静态与移动说话人场景。实验表明,相较于多个基线方法,所提音视频系统在静态与动态说话人场景下均展现出更优的语音增强性能与鲁棒性。
原文摘要 · Abstract (English)
Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal results in low signal-to-noise ratio (SNR) conditions, when there is high reverberation, or in complex scenarios involving dynamic speakers, overlapping speech, or non-stationary noise. To address these issues, we propose a novel Visual-Informed Neural Beamforming Network (VI-NBFNet), which integrates microphone array signal processing and deep neural networks (DNNs) using multimodal input features. The proposed network leverages a pretrained visual speech recognition model to extract lip movements as input features, which serve for voice activity detection (VAD) and target speaker identification. The system is intended to handle both static and moving speakers by introducing a supervised end-to-end beamforming framework equipped with an attention mechanism. The experimental results demonstrated that the proposed audiovisual system has achieved better SE performance and robustness for both stationary and dynamic speaker scenarios, compared to several baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。