arXiv:2505.20635eess.AScs.AI2025-05被引 1

让语音视频说话人分离更准,能同时处理多个在场人脸

Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction

  • 引入可插拔的跨说话人注意力机制,动态处理多个共现人脸
  • 在VoxCeleb2等数据集上指标提升,复杂场景表现更优
  • 适用于多人对话、会议等真实场景,兼容主流模型

音视频说话人提取旨在根据视觉线索从混合语音中分离出目标说话人语音,通常依赖目标说话人的面部视频。但在真实场景中,屏幕上常有多个共现人脸,这些信息可提供重要的说话人活动线索。本文提出一种即插即用的跨说话人注意力模块,用于处理任意数量的共现人脸,从而在复杂多说话人环境中实现更精确的说话人提取。我们将该模块集成到AV-DPRNN和当前最先进的AV-TFGridNet模型中。在多样化数据集上的大量实验表明,包括高度重叠的VoxCeleb2和稀疏重叠的MISP,本方法均持续优于基线模型。此外,在LRS2和LRS3上的跨数据集验证进一步证明了方法的鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co-occurring faces are often present on-screen, providing valuable speaker activity cues in the scene. In this work, we introduce a plug-and-play inter-speaker attention module to process these flexible numbers of co-occurring faces, allowing for more accurate speaker extraction in complex multi-person environments. We integrate our module into two prominent models: the AV-DPRNN and the state-of-the-art AV-TFGridNet. Extensive experiments on diverse datasets, including the highly overlapped VoxCeleb2 and sparsely overlapped MISP, demonstrate that our approach consistently outperforms baselines. Furthermore, cross-dataset evaluations on LRS2 and LRS3 confirm the robustness and generalizability of our method.

音视频融合说话人分离注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。