arXiv:2501.01518eess.AScs.SD2025-01CVPR被引 36

用视觉和文字信息提升嘈杂环境下的语音分离效果

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation

  • 融合视觉与文本信息,通过Transformer在原始波形上实现多模态语音分离
  • 在LRS2和LRS3数据集上达到当前最佳性能,且对视听不同步有强鲁棒性
  • 适合语音增强、智能会议系统等实际场景应用

本文旨在通过融合多种模态信息,在多说话人及噪声环境下实现语音分离与增强。已有研究证明,利用同步的唇部运动或人脸身份等时空视觉线索可取得良好效果。本文提出一种统一框架,支持同步或异步模态输入:(i)设计基于Transformer的现代架构,用于在原始波形域融合多模态信息;(ii)引入句子文本内容作为条件输入,或与视觉信息联合使用;(iii)验证模型对音频-视觉不同步偏移的鲁棒性;(iv)在知名基准数据集LRS2和LRS3上取得当前最优性能。

原文摘要 · Abstract (English)

The goal of this paper is speech separation and enhancement in multi-speaker and noisy environments using a combination of different modalities. Previous works have shown good performance when conditioning on temporal or static visual evidence such as synchronised lip movements or face identity. In this paper, we present a unified framework for multi-modal speech separation and enhancement based on synchronous or asynchronous cues. To that end we make the following contributions: (i) we design a modern Transformer-based architecture tailored to fuse different modalities to solve the speech separation task in the raw waveform domain; (ii) we propose conditioning on the textual content of a sentence alone or in combination with visual information; (iii) we demonstrate the robustness of our model to audio-visual synchronisation offsets; and, (iv) we obtain state-of-the-art performance on the well-established benchmark datasets LRS2 and LRS3.

语音分离多模态TransformerLRS3

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。