arXiv:2506.01270eess.AScs.SD2025-06被引 2

提出轻量级音视频联合语音分离模型,支持实时动态换主讲人场景。

Online Audio-Visual Autoregressive Speaker Extraction

  • 用深度可分离卷积设计轻量视觉前端,仅0.1万参数
  • 引入自回归声学编码器,提升分离效果0.9 dB(SI-SNRi)
  • 模型对注意力切换鲁棒,适合实际语音流场景

本文提出一种新型在线音视频说话人分离模型。在流式处理场景中,多数研究仅优化音频网络,忽视视觉前端的潜力。我们首先设计基于深度可分离卷积的轻量级视觉前端;随后提出轻量级自回归声学编码器,作为第二线索,主动利用前序步骤分离语音中的信息。首次在变化关注焦点(即目标说话人切换)的场景下评估算法表现。LRS3数据集实验表明,该视觉前端在SkiM与ConvTasNet音频骨干上均达到当前最优性能,仅需0.1百万网络参数和每秒2.1兆乘加运算(MACs)。自回归声学编码器进一步带来0.9 dB的SI-SNRi增益,且其动态特性对注意力变化具有鲁棒性。

原文摘要 · Abstract (English)

This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less explored. We first propose a lightweight visual frontend based on depth-wise separable convolution. Then, we propose a lightweight autoregressive acoustic encoder to serve as the second cue, to actively explore the information in the separated speech signal from past steps. Scenario-wise, for the first time, we study how the algorithm performs when there is a change in focus of attention, i.e., the target speaker. Experimental results on LRS3 datasets show that our visual frontend performs comparably to the previous state-of-the-art on both SkiM and ConvTasNet audio backbones with only 0.1 million network parameters and 2.1 MACs per second of processing. The autoregressive acoustic encoder provides an additional 0.9 dB gain in terms of SI-SNRi, and its momentum is robust against the change in attention.

语音分离音视频融合自回归轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。