提出MeMo框架,让语音分离系统在视觉信号缺失时仍能持续追踪目标说话人。
MeMo: Attentional Momentum for Real-Time Audio-Visual Target Speaker Extraction Under Impaired Visual Conditions
- 引入双自适应记忆库,保持注意力持续性
- 在视觉退化时仍实现至少2 dB的语音质量提升
- 适合实时语音分离场景,尤其视觉条件差时
音视频目标说话人分离(AV-TSE)通过视觉线索引导,从多人环境提取目标说话人语音。然而系统性能严重依赖视觉质量,在视觉线索缺失或严重退化时可能失效。人类即使无明确辅助信息也能持续关注目标说话人。受此启发,我们提出新框架MeMo,包含两个自适应记忆库以存储注意力相关信息。该框架专为实时场景设计:一旦建立初始注意力,系统便能在视觉线索不可用时维持注意力持续性。全面实验验证其有效性,结果表明,与基线相比,本方法在信噪比(SI-SNR)上至少提升2 dB。
原文摘要 · Abstract (English)
Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the performance of AV-TSE systems heavily relies on the quality of these visual cues. In extreme scenarios where visual cues are missing or severely degraded, the system may fail to accurately extract the target speaker. In contrast, humans can maintain attention on a target speaker even in the absence of explicit auxiliary information. Motivated by such human cognitive ability, we propose a novel framework called MeMo, which incorporates two adaptive memory banks to store attention-related information. MeMo is specifically designed for real-time scenarios: once initial attention is established, the system maintains attentional momentum over time, even when visual cues become unavailable. We conduct comprehensive experiments to verify the effectiveness of MeMo. Experimental results demonstrate that our proposed framework achieves SI-SNR improvements of at least 2 dB over the corresponding baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。