arXiv:2509.17741eess.AS2025-09被引 2

用生成模型提升多人环境下的目标语音分离效果

GAN-Based Multi-Microphone Spatial Target Speaker Extraction

  • 结合空间信息与判别模型中间特征,用GAN实现定向语音提取
  • 空间分辨率达5度,感知质量优于当前最优判别方法
  • 适合需要高精度语音分离的会议系统或智能音箱场景

空间目标说话人分离利用方向到达(DoA)等空间信息,在多人混叠环境中提取特定说话人的语音。尽管近年来基于深度神经网络(DNN)的判别方法已取得显著性能提升,但生成式方法如生成对抗网络(GAN)在此任务中的潜力尚未被充分探索。本文证明,通过将噪声混合信号与空间信息联合输入,一个条件GAN可有效提取并重建目标说话人的语音。具体地,我们在训练中引入判别式空间滤波模型的中间特征,并以DoA为条件,实现可调控的目标语音提取,空间分辨率达5度,在基于感知质量的客观指标上超越现有最优判别方法。

原文摘要 · Abstract (English)

Spatial target speaker extraction isolates a desired speaker's voice in multi-speaker environments using spatial information, such as the direction of arrival (DoA). Although recent deep neural network (DNN)-based discriminative methods have shown significant performance improvements, the potential of generative approaches, such as generative adversarial networks (GANs), remains largely unexplored for this problem. In this work, we demonstrate that a GAN can effectively leverage both noisy mixtures and spatial information to extract and generate the target speaker's speech. By conditioning the GAN on intermediate features of a discriminative spatial filtering model in addition to DoA, we enable steerable target extraction with high spatial resolution of 5 degrees, outperforming state-of-the-art discriminative methods in perceptual quality-based objective metrics.

语音分离GAN空间信息多麦克风

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。