用少量麦克风生成虚拟信号,提升语音增强效果。
Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement

- 通过神经网络从有限麦克风数据生成虚拟麦克风信号。
- 在多种语音分离系统中接近全麦克风配置的性能。
- 适合资源受限设备上的高精度语音增强应用。
多麦克风语音增强算法的定向性能随麦克风数量增加而提升,但实际边缘设备受限于物理空间难以部署大型阵列。为克服此限制,本文提出 Spatial-Magnifier,一种神经网络模型,可从有限的真实麦克风(RM)测量值生成虚拟麦克风(VM)信号。同时引入 Spatial Audio Representation Learning(SARL)框架,利用估计的虚拟麦克风信号与特征来指导下游语音增强系统。实验表明,该框架在多种语音提取系统中均优于现有空间上采样基线,包括端到端多通道语音增强和神经波束成形。所提方法几乎恢复了所有麦克风可用时的最优性能(oracle performance)。
原文摘要 · Abstract (English)
While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this limitation, we propose Spatial-Magnifier, a neural network designed to generate virtual microphone (VM) signals from a limited set of real microphone (RM) measurements. Moreover, we introduce the Spatial Audio Representation Learning (SARL) framework, which leverages estimated VM signals and features to condition a downstream speech enhancement system. Experimental results demonstrate that the proposed framework outperforms existing spatial upsampling baselines across various speech extraction systems, including end-to-end multichannel speech enhancement and neural beamforming. The proposed method nearly recovers the oracle performance achieved when all microphones are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。