arXiv:2601.19130eess.AS2026-01

融合手势与嘴部动作,提升语音分离鲁棒性

Beyond Lips: Integrating Gesture and Lip Cues for Robust Audio-visual Speaker Extraction

  • 用跨注意力机制融合嘴部与上身手势信息
  • 对比学习使手势表征更贴近语音,提升分离效果
  • 在遮挡或远距离时仍表现优异,适合真实场景

多数音视频说话人分离方法依赖同步的嘴部影像来分离目标说话人语音。但在自然交流中,伴随手势也与语音同步,常强调特定词或音节,提供补充视觉线索,尤其在面部或嘴部被遮挡或距离较远时更为重要。本文提出SeLG模型,突破以嘴部为中心的方法,融合嘴部与上身手势信息实现鲁棒说话人分离。SeLG采用基于交叉注意力的融合机制,使每种视觉模态可查询并选择性关注混合语音中的相关特征。为增强手势表征与语音动态对齐,还引入对比损失(InfoNCE),促使手势嵌入更贴近与语音强相关的嘴部嵌入。在包含TED演讲的YGD数据集上的实验表明,该对比学习策略显著提升了基于手势的说话人分离性能;且在完整与部分(即缺失模态)条件下,所提模型均优于基线,展现出更强鲁棒性。

原文摘要 · Abstract (English)

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned with speech, often emphasizing specific words or syllables. These gestures provide complementary visual cues that can be especially valuable when facial or lip regions are occluded or distant. In this work, we move beyond lip-centric approaches and propose SeLG, a model that integrates both lip and upper-body gesture information for robust speaker extraction. SeLG features a cross-attention-based fusion mechanism that enables each visual modality to query and selectively attend to relevant speech features in the mixture. To improve the alignment of gesture representations with speech dynamics, SeLG also employs a contrastive InfoNCE loss that encourages gesture embeddings to align more closely with corresponding lip embeddings, which are more strongly correlated with speech. Experimental results on the YGD dataset, containing TED talks, demonstrate that the proposed contrastive learning strategy significantly improves gesture-based speaker extraction, and that our proposed SeLG model, by effectively fusing lip and gesture cues with an attention mechanism and InfoNCE loss, achieves superior performance compared to baselines, across both complete and partial (i.e., missing-modality) conditions.

音视频分离多模态融合手势识别鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。