arXiv:2506.06689cs.SDeess.AS2025-06中稿 · ECAI 2025被引 3

轻量级模型实现实时音视频语音分离,支持低延迟处理。

A Fast and Lightweight Model for Causal Audio-Visual Speech Separation

  • 采用轻量化视觉模块与高效融合机制,支持流式处理。
  • 在LRS2、LRS3、VoxCeleb2上表现优于现有因果方法。
  • 可将非因果模型转为因果模型,适合实时应用场景。

音视频语音分离(AVSS)通过结合听觉与视觉(唇动)信息,从混合信号中提取目标语音。然而,多数现有方法架构复杂且依赖未来上下文,需离线处理,难以用于实时应用。受RTFSNet启发,本文提出一种新型流式AVSS模型Swift-Net,增强实时处理所需的因果性。Swift-Net采用轻量级视觉特征提取模块与高效音频-视觉融合模块,并利用分组SRU在不同特征空间中整合历史信息,提升历史信息利用效率。此外,提出因果转换模板,可将非因果AVSS模型转化为因果版本。在三个标准数据集(LRS2、LRS3、VoxCeleb2)上的实验表明,在因果条件下,Swift-Net展现出优异性能,凸显其在复杂环境下的语音处理潜力。

原文摘要 · Abstract (English)

Audio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing AVSS methods exhibit complex architectures and rely on future context, operating offline, which renders them unsuitable for real-time applications. Inspired by the pipeline of RTFSNet, we propose a novel streaming AVSS model, named Swift-Net, which enhances the causal processing capabilities required for real-time applications. Swift-Net adopts a lightweight visual feature extraction module and an efficient fusion module for audio-visual integration. Additionally, Swift-Net employs Grouped SRUs to integrate historical information across different feature spaces, thereby improving the utilization efficiency of historical information. We further propose a causal transformation template to facilitate the conversion of non-causal AVSS models into causal counterparts. Experiments on three standard benchmark datasets (LRS2, LRS3, and VoxCeleb2) demonstrated that under causal conditions, our proposed Swift-Net exhibited outstanding performance, highlighting the potential of this method for processing speech in complex environments.

语音分离音视频实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。