用扩散模型提升少通道麦克风阵列的声源方向分辨率
SIRUP: A diffusion-based virtual upmixer of steering vectors for highly-directive spatialization with first-order ambisonics
- 用变分自编码器与扩散模型联合学习高阶音场嵌入
- 在真实数据上实现比传统方法更高的声源定位精度
- 适合做空间音频增强和语音降噪的研究者
本文提出一种基于潜在扩散模型的虚拟升维方法,用于从少通道球形麦克风阵列捕获的首阶全向声学(FOA)数据中重建高阶音场(HOA)的导向矢量。传统方法依赖物理声学模拟器,但受限于FOA的空间分辨率与声源方向估计之间的耦合关系。SIRUP采用变分自编码器(VAE)在隐空间压缩HOA数据,再训练扩散模型以FOA为条件生成对应的高阶嵌入。实验表明,该方法在导向矢量升维、声源定位及语音降噪任务上均显著优于传统FOA系统。
原文摘要 · Abstract (English)
This paper presents virtual upmixing of steering vectors captured by a fewer-channel spherical microphone array. This challenge has conventionally been addressed by recovering the directions and signals of sound sources from first-order ambisonics (FOA) data, and then rendering the higher-order ambisonics (HOA) data using a physics-based acoustic simulator. This approach, however, struggles to handle the mutual dependency between the spatial directivity of source estimation and the spatial resolution of FOA ambisonics data. Our method, named SIRUP, employs a latent diffusion model architecture. Specifically, a variational autoencoder (VAE) is used to learn a compact encoding of the HOA data in a latent space and a diffusion model is then trained to generate the HOA embeddings, conditioned by the FOA data. Experimental results showed that SIRUP achieved a significant improvement compared to FOA systems for steering vector upmixing, source localization, and speech denoising.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。