用图神经网络建模声音空间相关性,快速生成个性化耳机音频数据。
Graph Neural Field with Spatial-Correlation Augmentation for HRTF Personalization
- 通过图网络预测目标用户的声学特征,结合编码器-解码器结构。
- 引入空间相关性增强模块,提升不同位置声音的一致性表现。
- 适合虚拟现实/增强现实中的个性化音频渲染,无需逐点测量。
为实现虚拟现实/增强现实设备上的沉浸式空间音频渲染,高质量的头相关传输函数(HRTF)至关重要。通常,HRTF具有个体和位置依赖性,测量过程耗时且繁琐。为此,本文提出图神经场与空间相关性增强方法(GraphNF-SCA),用于生成未见个体的个性化HRTF。该方法包含三个核心模块:HRTF个性化(HRTF-P)模块、HRTF上采样(HRTF-U)模块和微调阶段。HRTF-P模块采用编码器-解码器架构的图神经网络,编码器提取通用特征,解码器融合目标相关特征并输出个性化HRTF。HRTF-U模块使用另一个图神经网络建模不同位置间HRTF的空间相关性,通过微调提升预测结果的空间一致性。相比以往仅逐点估计而忽略空间关联的方法,GraphNF-SCA有效利用了HRTF内在的空间相关性,显著提升了个性化性能。实验表明,该方法达到当前最优效果。
原文摘要 · Abstract (English)
To achieve immersive spatial audio rendering on VR/AR devices, high-quality Head-Related Transfer Functions (HRTFs) are essential. In general, HRTFs are subject-dependent and position-dependent, and their measurement is time-consuming and tedious. To address this challenge, we propose the Graph Neural Field with Spatial-Correlation Augmentation (GraphNF-SCA) for HRTF personalization, which can be used to generate individual HRTFs for unseen subjects. The GraphNF-SCA consists of three key components: an HRTF personalization (HRTF-P) module, an HRTF upsampling (HRTF-U) module, and a fine-tuning stage. In the HRTF-P module, we predict HRTFs of the target subject via the Graph Neural Network (GNN) with an encoder-decoder architecture, where the encoder extracts universal features and the decoder incorporates the target-relevant features and produces individualized HRTFs. The HRTF-U module employs another GNN to model spatial correlations across HRTFs. This module is fine-tuned using the output of the HRTF-P module, thereby enhancing the spatial consistency of the predicted HRTFs. Unlike existing methods that estimate individual HRTFs position-by-position without spatial correlation modeling, the GraphNF-SCA effectively leverages inherent spatial correlations across HRTFs to enhance the performance of HRTF personalization. Experimental results demonstrate that the GraphNF-SCA achieves state-of-the-art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。