arXiv:2501.01401eess.AS2025-01被引 3

用音视频联合建模,实现精准目标人声分离。

VoiceVector: Multimodal Enrolment Vectors for Speaker Separation

  • 通过音频与视觉信息构建说话人专属嵌入向量
  • 支持仅用视频唇动或纯音频生成嵌入向量
  • 可同时使用正负样本增强分离效果,适合多说话人场景

我们提出一种基于Transformer的语音分离架构,从多个说话人和背景噪声中分离目标说话人。系统包含两个独立神经网络:(A) 预注册网络,利用音频、视觉(如口型运动)或两者组合生成说话人特异性嵌入向量;(B) 分离网络,接收含噪信号与注册向量,输出目标说话人的纯净语音。创新点在于:(i) 注册向量可由纯音频、音视频数据或仅无声视频中的唇动生成;(ii) 可灵活使用多个正负注册向量进行条件化分离。相比先前方法,性能显著提升。

原文摘要 · Abstract (English)

We present a transformer-based architecture for voice separation of a target speaker from multiple other speakers and ambient noise. We achieve this by using two separate neural networks: (A) An enrolment network designed to craft speaker-specific embeddings, exploiting various combinations of audio and visual modalities; and (B) A separation network that accepts both the noisy signal and enrolment vectors as inputs, outputting the clean signal of the target speaker. The novelties are: (i) the enrolment vector can be produced from: audio only, audio-visual data (using lip movements) or visual data alone (using lip movements from silent video); and (ii) the flexibility in conditioning the separation on multiple positive and negative enrolment vectors. We compare with previous methods and obtain superior performance.

说话人分离多模态音视频融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。