arXiv:2507.21448eess.AScs.ET2025-07中稿 · to Interspeech 202…被引 3

用视觉信息实时增强语音,有效分离目标说话人。

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

  • 融合视听语音识别与说话人定位的视觉嵌入
  • 低信噪比多说话人场景下性能提升显著
  • 首个开源实时音频视觉语音增强系统

在仅有音频输入的情况下,抑制干扰说话人仍具挑战。本文提出一种简单高效的实时音视频语音增强(AVSE)系统RAVEN,可分离并增强屏幕上的目标说话人,同时抑制干扰说话人和背景噪声。我们研究了来自视听语音识别(AVSR)和主动说话人检测(ASD)模型学习的视觉嵌入在不同信噪比(SNR)条件及干扰说话人数下的作用。结果表明,在低信噪比、多说话人环境下,融合AVSR与ASD嵌入效果最佳;而在仅存在背景噪声时,仅使用AVSR嵌入表现最优。此外,我们开发了一个可在计算机CPU上实时运行的流式系统,并提供视频演示与代码仓库。据我们所知,这是首个开源的实时音视频语音增强系统。

原文摘要 · Abstract (English)

Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAVEN, which isolates and enhances the on-screen target speaker while suppressing interfering speakers and background noise. We investigate how visual embeddings learned from audio-visual speech recognition (AVSR) and active speaker detection (ASD) contribute to AVSE across different SNR conditions and numbers of interfering speakers. Our results show concatenating embeddings from AVSR and ASD models provides the greatest improvement in low-SNR, multi-speaker environments, while AVSR embeddings alone perform best in noise-only scenarios. In addition, we develop a real-time streaming system that operates on a computer CPU and we provide a video demonstration and code repository. To our knowledge, this is the first open-source implementation of a real-time AVSE system.

语音增强视觉融合实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。