arXiv:2409.02266cs.SDcs.LG2024-09被引 11

用视听融合提升语音降噪,性能超越基准模型。

LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement

  • 融合视觉与音频特征,通过时序建模增强语音信号。
  • 在SISDR、STOI、PESQ上分别提升0.06、0.03、1.32。
  • 适合研究多模态语音增强或竞赛复现的开发者。

本文提出一种长短期记忆语音增强网络(LSTMSE-Net),用于视听语音增强(AVSE)。该方法利用视觉与音频信息的互补性,提升语音质量。视觉特征由VisualFeatNet(VFN)提取,音频特征经编码器-解码器处理。系统对视听特征进行缩放与拼接,再通过分离网络实现优化增强。该架构在多模态数据融合与插值技术方面取得进展,显著提升鲁棒性。LSTMSE-Net在COG-MHEAR AVSE挑战赛2024基线模型基础上,实现0.06的尺度无关信干比(SISDR)提升、0.03的短时客观可懂度(STOI)提升,以及1.32的语音质量感知评价(PESQ)提升。源代码已公开于https://github.com/mtanveer1/AVSEC-3-Challenge。

原文摘要 · Abstract (English)

In this paper, we propose long short term memory speech enhancement network (LSTMSE-Net), an audio-visual speech enhancement (AVSE) method. This innovative method leverages the complementary nature of visual and audio information to boost the quality of speech signals. Visual features are extracted with VisualFeatNet (VFN), and audio features are processed through an encoder and decoder. The system scales and concatenates visual and audio features, then processes them through a separator network for optimized speech enhancement. The architecture highlights advancements in leveraging multi-modal data and interpolation techniques for robust AVSE challenge systems. The performance of LSTMSE-Net surpasses that of the baseline model from the COG-MHEAR AVSE Challenge 2024 by a margin of 0.06 in scale-invariant signal-to-distortion ratio (SISDR), $0.03$ in short-time objective intelligibility (STOI), and $1.32$ in perceptual evaluation of speech quality (PESQ). The source code of the proposed LSTMSE-Net is available at \url{https://github.com/mtanveer1/AVSEC-3-Challenge}.

视听融合语音增强深度学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。