arXiv:2603.07696eess.AS2026-03被引 1

用多视角视频提升语音分离效果,让单视角也能受益。

Multi-View Based Audio Visual Target Speaker Extraction

  • 通过多视角唇部特征的张量融合,建模跨视角交互关系。
  • 单视角输入下性能显著提升,多视角输入时更稳健。
  • 适合需要鲁棒语音分离的实际场景,如会议、监控等。

音频-视觉目标说话人分离(AVTSE)旨在利用对应视觉线索从混叠音频中分离出目标说话人的声音。现有方法大多依赖正面视角视频,限制了在真实场景中的鲁棒性,而侧视等非正面视角常包含互补的发音信息。本文提出多视角张量融合(MVTF)框架,将多视角学习转化为单视角性能提升。训练阶段,利用同步的多角度唇部视频,通过成对外积显式建模不同视角输入唇部嵌入间的乘积交互关系。推理阶段,系统支持单视角与多视角输入。实验表明,单视角输入时,模型借助多视角知识实现显著性能提升;多视角输入时,整体性能进一步优化并增强鲁棒性。演示、代码与数据可在 https://anonymous.4open.science/w/MVTF-Gridnet-209C/ 获取。

原文摘要 · Abstract (English)

Audio-Visual Target Speaker Extraction (AVTSE) aims to separate a target speaker's voice from a mixed audio signal using the corresponding visual cues. While most existing AVTSE methods rely exclusively on frontal-view videos, this limitation restricts their robustness in real-world scenarios where non-frontal views are prevalent. Such visual perspectives often contain complementary articulatory information that could enhance speech extraction. In this work, we propose Multi-View Tensor Fusion (MVTF), a novel framework that transforms multi-view learning into single-view performance gains. During the training stage, we leverage synchronized multi-perspective lip videos to learn cross-view correlations through MVTF, where pairwise outer products explicitly model multiplicative interactions between different views of input lip embeddings. At the inference stage, the system supports both single-view and multi-view inputs. Experimental results show that in the single-view inputs, our framework leverages multi-view knowledge to achieve significant performance gains, while in the multi-view mode, it further improves overall performance and enhances the robustness. Our demo, code and data are available at https://anonymous.4open.science/w/MVTF-Gridnet-209C/

语音分离多视角视听融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。