arXiv:2509.12583eess.AScs.SD2025-09

融合单帧人脸与帧级唇动特征,提升语音提取鲁棒性

Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion

  • 用单帧人脸与帧级唇动互补融合,增强多模态感知
  • 高缺失率训练使模型在测试时丢失模态仍保持稳定性能
  • 适合真实场景中信号不全的语音提取任务

音频-视觉目标说话人提取(AVTSE)在嘈杂聚会场景中至关重要。利用多源线索——如话语级说话人嵌入或静态人脸图像,以及帧级唇部运动或面部表情特征——可显著提升性能。然而,实际应用中常出现帧级线索间歇性丢失。本文系统研究了多注册融合在不同模态缺失情况下的鲁棒性。结果表明,尽管全模态融合在理想条件下表现优异,但在测试时遇到未见模态缺失时性能急剧下降。关键发现是:以高缺失率进行训练能显著提升鲁棒性,即使在严重测试时模态缺失下仍保持稳定表现。实验验证,将单帧人脸图像与帧级唇动特征融合,可在保证强性能的同时实现高鲁棒性。模型与代码已公开。

原文摘要 · Abstract (English)

Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images, and frame-level lip motion or facial expression features --can significantly improve performance. However, real-world applications often suffer from intermittent signal loss, especially for frame-level cues. This paper systematically investigates the robustness of multi-enrollment fusion under varying degrees of modality missing. Results show that while full multimodal fusion excels under ideal conditions, its performance degrades sharply when encountering unseen modalities missing during the testing. Crucially, training with a high missing rate dramatically enhances robustness, maintaining stable performance even under severe test-time modality missing. We demonstrate that fusing the complementary one frame of face image with frame-level lip features achieves both strong performance and robustness for the AVTSE task. The model and codes are shared.

语音提取多模态融合鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。