先分声源再匹配,提升真实场景下的音视频语音增强效果。
Separate First, Then Associate: A Two-Stage Approach for Real-World Audio-Visual Speech Enhancement

- 分两阶段:先用纯音频模型分离多人说话声,再用音视频CLIP匹配目标说话人
- 在真实数据集上表现优于直接端到端方法,有效应对重叠、混响等复杂情况
- 适合需要高鲁棒性的实际音视频应用,如会议记录、监控分析
音视频语音增强(AVSE)旨在利用视觉线索从多人混合语音中提取目标语音。尽管近期研究在模拟数据集上表现优异,但在真实录音场景下性能往往大幅下降。为弥合这一差距,ISCSLP 2026会议举办了真实世界AVSE挑战赛,要求设计能在真实条件下处理说话人重叠、声学干扰、房间混响及视觉退化的实用解决方案。本文提交方案采用解耦的分离-关联两阶段方法:第一阶段使用训练好的纯音频模型将多说话人混合信号分离为单个说话人信号;第二阶段则利用音视频CLIP模型,通过跨模态相似性匹配,识别与目标说话人面部视频最相似的分离语音信号。在挑战赛数据集上的评估结果验证了该方法的有效性。
原文摘要 · Abstract (English)
Audio-visual speech enhancement (AVSE) aims at extracting target speech from multi-speaker mixtures by exploiting visual cues. Although recent studies have reported strong performance on simulated datasets, the performance, however, often drops dramatically when they are applied to real-world audio-visual recordings. To bridge this gap, the Real-World AVSE Challenge held in the ISCSLP 2026 conference calls for participants to design a practical solution for AVSE under real-world conditions, where speaker overlap, acoustic interferences, room reverberation and visual degradations naturally co-exist. In our submission to the challenge, we propose a decoupled separation-then-association approach. It consists of two stages: a separation stage in which a trained, audio-only model (i.e., not using visual cues) is used to separate input multi-speaker mixture to individual speaker signals, followed by an association stage, where an audio-visual CLIP model is used to identify the separated speech signal with the highest similarity with the target speaker's facial video via cross-modal similarity matching. Evaluation results on the challenge dataset show the effectiveness of our proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。