arXiv:2509.22425cs.SD2025-09

通过递归增强视觉语义,提升语音分离精度。

From Coarse to Fine: Recursive Audio-Visual Semantic Enhancement for Speech Separation

  • 分两阶段:先粗分离再用视觉反馈精修音频
  • 在三个基准数据集上达到当前最佳性能
  • 适合需要高精度语音分离的研究者

音视频语音分离旨在利用唇动等视觉线索从混合语音中分离出每个说话人的清晰语音。现有方法多依赖静态视觉表示,未能充分挖掘视觉信息潜力。本文提出CSFNet——一种从粗到精的递归语义增强网络。该模型分两阶段运行:(1)粗分离阶段,基于混合音频和视觉输入重建初步音频波形;(2)精分离阶段,将粗分离结果与视觉流共同输入音视频语音识别模型,实现递归反馈。此过程生成更具判别性的语义表征,用于提取优化后的音频。为更好利用这些语义,设计了说话人感知的感知融合模块以跨模态编码身份信息,并引入多尺度时频分离网络捕捉局部与全局时间-频率模式。在三个基准数据集及两个噪声数据集上的大量实验表明,CSFNet取得当前最优性能,且粗到精提升显著,验证了递归语义增强框架的有效性。

原文摘要 · Abstract (English)

Audio-visual speech separation aims to isolate each speaker's clean voice from mixtures by leveraging visual cues such as lip movements and facial features. While visual information provides complementary semantic guidance, existing methods often underexploit its potential by relying on static visual representations. In this paper, we propose CSFNet, a Coarse-to-Separate-Fine Network that introduces a recursive semantic enhancement paradigm for more effective separation. CSFNet operates in two stages: (1) Coarse Separation, where a first-pass estimation reconstructs a coarse audio waveform from the mixture and visual input; and (2) Fine Separation, where the coarse audio is fed back into an audio-visual speech recognition (AVSR) model together with the visual stream. This recursive process produces more discriminative semantic representations, which are then used to extract refined audio. To further exploit these semantics, we design a speaker-aware perceptual fusion block to encode speaker identity across modalities, and a multi-range spectro-temporal separation network to capture both local and global time-frequency patterns. Extensive experiments on three benchmark datasets and two noisy datasets show that CSFNet achieves state-of-the-art (SOTA) performance, with substantial coarse-to-fine improvements, validating the necessity and effectiveness of our recursive semantic enhancement framework.

语音分离音视频融合递归增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。