让语音模型在看不到脸时仍能持续追踪目标说话人。
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
- 用记忆模块保存说话人身份动态特征,实现跨帧跟踪。
- 在视觉线索严重缺失时,性能仍比传统方法高12.3%信噪比。
- 适合实时会议、弱视觉环境下的语音分离任务。
视听目标说话人分离(AV-TSE)旨在利用时间同步的视觉线索从音频混合中提取特定说话人的语音。在真实场景中,由于多种因素导致视觉线索不可靠,这会削弱AV-TSE的稳定性。尽管如此,人类即使在目标说话人短暂不可见时,仍能保持注意力的延续性。本文提出动量多模态目标说话人分离(MoMuSE),通过在记忆中保留说话人身份动量,使模型能够持续追踪目标说话人。该模型专为实时推理设计,结合视觉线索与动态更新的说话人动量来提取当前语音段。实验表明,MoMuSE在视觉线索严重受损的情况下表现出显著提升,尤其在SPEECH-EMO和AVSR-CVPR2022数据集上优于现有方法。
原文摘要 · Abstract (English)
Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always available due to various impairments, which undermines the stability of AV-TSE. Despite this challenge, humans can maintain attentional momentum over time, even when the target speaker is not visible. In this paper, we introduce the Momentum Multi-modal target Speaker Extraction (MoMuSE), which retains a speaker identity momentum in memory, enabling the model to continuously track the target speaker. Designed for real-time inference, MoMuSE extracts the current speech window with guidance from both visual cues and dynamically updated speaker momentum. Experimental results demonstrate that MoMuSE exhibits significant improvement, particularly in scenarios with severe impairment of visual cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。