arXiv:2509.20741eess.AScs.ET2025-09中稿 · to WASPAA 2025 dem…

用嘴型信息实时增强语音,纯CPU运行,效果稳定。

Real-Time System for Audio-Visual Target Speech Enhancement

  • 结合嘴型视觉信号,提升嘈杂环境下的语音清晰度。
  • 支持单麦克风、多干扰源、甚至唱歌场景的实时处理。
  • 无需专用硬件,适合普通电脑部署,体验真实互动。

我们展示了RAVEN系统,一个可在纯CPU上实时运行的音视频目标语音增强系统。传统单通道语音增强主要针对环境噪声,近年研究开始利用嘴型等视觉线索提升鲁棒性,尤其在存在干扰说话者时。然而,目前尚无公开演示可实现基于CPU的实时音视频语音增强。RAVEN通过预训练音视频语音识别模型提取的视觉嵌入,编码唇动信息,有效泛化于各类噪声、干扰人声、突发声音及歌唱场景。本演示中,用户可通过麦克风和摄像头实时体验语音增强,并通过耳机收听纯净语音输出。

原文摘要 · Abstract (English)

We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as the task of extracting clean speech from environmental noise. More recent work has explored the use of visual cues, such as lip movements, to improve robustness, particularly in the presence of interfering speakers. However, to our knowledge, no prior work has demonstrated an interactive system for real-time audio-visual speech enhancement operating on CPU hardware. RAVEN fills this gap by using pretrained visual embeddings from an audio-visual speech recognition model to encode lip movement information. The system generalizes across environmental noise, interfering speakers, transient sounds, and even singing voices. In this demonstration, attendees will be able to experience live audio-visual target speech enhancement using a microphone and webcam setup, with clean speech playback through headphones.

语音增强音视频融合实时系统纯CPU

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。