arXiv:2508.13624cs.SDeess.AS2025-08中稿 · Interspeech 2025 W…被引 1

用全脸视觉信息增强语音,让模型在嘈杂环境里听得更准。

Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement

  • 结合全脸视觉与Mamba结构,捕捉时空视觉线索提升语音分离能力。
  • 在音频视觉增强挑战赛中,各项指标领先,单声道榜排名第一。
  • 适合做复杂场景下的语音增强,如多人会议或嘈杂环境应用。

近期基于Mamba的模型在语音增强任务中展现出高效建模长时依赖的潜力。然而,诸如语音增强Mamba(SEMamba)等模型仍局限于单说话人场景,在复杂的多说话人环境(如鸡尾酒会问题)中表现不佳。为此,我们提出AVSEMamba,一种融合全脸视觉线索与Mamba时间骨干网络的音视频语音增强模型。通过利用时空视觉信息,AVSEMamba显著提升了复杂条件下的目标语音提取精度。在AVSEC-4挑战赛的开发集和盲测集上评估,该模型在语音可懂度(STOI)、感知质量(PESQ)及非侵入式质量(UTMOS)三项指标上均优于其他单声道基线模型,并在单声道排行榜上取得第一的成绩。

原文摘要 · Abstract (English)

Recent Mamba-based models have shown promise in speech enhancement by efficiently modeling long-range temporal dependencies. However, models like Speech Enhancement Mamba (SEMamba) remain limited to single-speaker scenarios and struggle in complex multi-speaker environments such as the cocktail party problem. To overcome this, we introduce AVSEMamba, an audio-visual speech enhancement model that integrates full-face visual cues with a Mamba-based temporal backbone. By leveraging spatiotemporal visual information, AVSEMamba enables more accurate extraction of target speech in challenging conditions. Evaluated on the AVSEC-4 Challenge development and blind test sets, AVSEMamba outperforms other monaural baselines in speech intelligibility (STOI), perceptual quality (PESQ), and non-intrusive quality (UTMOS), and achieves \textbf{1st place} on the monaural leaderboard.

语音增强音视频融合Mamba多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。