arXiv:2510.16437eess.AS2025-10

融合视觉与空间音频信息,提升嘈杂环境中的语音可懂度。

Audio-Visual Speech Enhancement for Spatial Audio - Spatial-VisualVoice and the MAVE Database

  • 用视觉和麦克风阵列的空间线索联合增强目标语音。
  • 低信噪比下显著提升语音质量,指标优于传统方法。
  • 适合虚拟现实、会议系统等需要空间感知的场景。

音频-视觉语音增强(AVSE)在低信噪比(SNR)条件下表现优异,因视觉特征对声学噪声具有鲁棒性。然而,针对低SNR下空间音频增强的AVSE方法仍存在明显空白,这在增强现实应用中日益重要。为此,我们提出一种基于VisualVoice的多通道AVSE框架,利用麦克风阵列的空间线索与视觉信息,在嘈杂环境中增强目标说话人语音。同时,我们构建了MAVe数据库,包含在可控、可复现的房间条件下采集的多通道音视频信号,覆盖广泛SNR范围。实验表明,该方法在低SNR下持续显著提升SI-SDR、STOI和PESQ指标。双耳信号分析进一步验证了空间线索与语音可懂度的有效保留。

原文摘要 · Abstract (English)

Audio-visual speech enhancement (AVSE) has been found to be particularly useful at low signal-to-noise (SNR) ratios due to the immunity of the visual features to acoustic noise. However, a significant gap exists in AVSE methods tailored to enhance spatial audio under low-SNR conditions. The latter is of growing interest with augmented reality applications. To address this gap, we present a multi-channel AVSE framework based on VisualVoice that leverages spatial cues from microphone arrays and visual information for enhancing the target speaker in noisy environments. We also introduce MAVe, a novel database containing multi-channel audio-visual signals in controlled, reproducible room conditions across a wide range of SNR levels. Experiments demonstrate that the proposed method consistently achieves significant gains in SI-SDR, STOI, and PESQ, particularly in low SNRs. Binaural signal analysis further confirms the preservation of spatial cues and intelligibility.

语音增强空间音频视听融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。