arXiv:2510.08914cs.SDeess.AS2025-10

通过虚拟麦克风提升低麦克风场景下的语音分离效果

VM-UNSSOR: Unsupervised Neural Speech Separation Enhanced by Higher-SNR Virtual Microphone Arrays

  • 用线性空间解混生成高信噪比虚拟麦克风信号
  • 在6麦克风下达到17.1 dB SI-SDR,2麦克风下达10.7 dB
  • 适合麦克风数量受限的无监督语音分离任务

盲语音分离(BSS)旨在未知阵列几何和房间冲激响应条件下,从多通道、多说话人混合信号中恢复多个语音源。在无监督设置中,由于缺乏干净目标语音用于模型训练,UNSSOR提出混合一致性(MC)损失,使深度神经网络在过定混合信号上训练以实现无监督语音分离。然而,当训练混合信号的麦克风数量减少时,MC约束减弱,分离性能急剧下降。为此,我们提出VM-UNSSOR,通过将有限麦克风记录的观测混合信号,经由线性空间解混器(如IVA和空间聚类)处理,生成若干高信噪比的虚拟麦克风(VM)信号,增强训练数据。作为观测混合信号的线性投影,虚拟麦克风信号通常能提升各源的信噪比,并可用于计算额外的MC损失,从而改进UNSSOR并缓解其频率排列问题。在SMS-WSJ数据集上,6麦克风、双说话人设置下,VM-UNSSOR达到17.1 dB SI-SDR,而UNSSOR仅14.7 dB;在确定的2麦克风、双说话人情况下,UNSSOR性能崩溃至-2.7 dB SI-SDR,而VM-UNSSOR达到10.7 dB。

原文摘要 · Abstract (English)

Blind speech separation (BSS) aims to recover multiple speech sources from multi-channel, multi-speaker mixtures under unknown array geometry and room impulse responses. In unsupervised setup where clean target speech is not available for model training, UNSSOR proposes a mixture consistency (MC) loss for training deep neural networks (DNN) on over-determined training mixtures to realize unsupervised speech separation. However, when the number of microphones of the training mixtures decreases, the MC constraint weakens and the separation performance falls dramatically. To address this, we propose VM-UNSSOR, augmenting the observed training mixture signals recorded by a limited number of microphones with several higher-SNR virtual-microphone (VM) signals, which are obtained by applying linear spatial demixers (such as IVA and spatial clustering) to the observed training mixtures. As linear projections of the observed mixtures, the virtual-microphone signals can typically increase the SNR of each source and can be leveraged to compute extra MC losses to improve UNSSOR and address the frequency permutation problem in UNSSOR. On the SMS-WSJ dataset, in the over-determined six-microphone, two-speaker separation setup, VM-UNSSOR reaches 17.1 dB SI-SDR, while UNSSOR only obtains 14.7 dB; and in the determined two-microphone, two-speaker case, UNSSOR collapses to -2.7 dB SI-SDR, while VM-UNSSOR achieves 10.7 dB.

语音分离无监督学习虚拟麦克风语音增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。