用人体参数预测个人听觉特征,提升虚拟音频真实感。
Head-Related Transfer Function Individualization Using Anthropometric Features and Spatially Independent Latent Representation
- 通过自编码器提取声源位置独立的听觉表征,降低模型复杂度。
- 在有限数据下实现高精度听觉特征预测,优于现有深度学习方法。
- 适合音频渲染、虚拟现实等需个性化听觉体验的场景。
提出一种基于受试者人体测量参数进行头相关传输函数(HRTF)个体化的方法。由于测量成本高,许多HRTF数据集的受试者数量有限,而包含人体测量参数的数据更少,因此基于深度神经网络(DNN)的HRTF个体化仍具挑战性。本文提出一种方法,利用自编码器对不同声源位置下的HRTF幅度进行条件建模,获得空间无关的潜在表征,从而可融合多个具有不同声源位置的HRTF数据集,并通过减少由人体参数估计的参数量,使网络训练更可行。实验表明,该方法在预测精度上优于当前主流的基于DNN的方法。
原文摘要 · Abstract (English)
A method for head-related transfer function (HRTF) individualization from the subject's anthropometric parameters is proposed. Due to the high cost of measurement, the number of subjects included in many HRTF datasets is limited, and the number of those that include anthropometric parameters is even smaller. Therefore, HRTF individualization based on deep neural networks (DNNs) is a challenging task. We propose a HRTF individualization method using the latent representation of HRTF magnitude obtained through an autoencoder conditioned on sound source positions, which makes it possible to combine multiple HRTF datasets with different measured source positions, and makes the network training tractable by reducing the number of parameters to be estimated from anthropometric parameters. Experimental evaluation shows that high estimation accuracy is achieved by the proposed method, compared to current DNN-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。