arXiv:2501.02871cs.SDeess.AS2025-01被引 8

用扩散模型根据人体数据生成个性化声学响应,提升沉浸音频真实感。

Towards HRTF Personalization using Denoising Diffusion Models

  • 以人体测量数据为条件,用扩散模型生成个性化头相关脉冲响应。
  • 生成效果达到当前最优水平,验证了方法可行性。
  • 适合做沉浸式音频、虚拟现实的开发者与研究者参考。

头相关传输函数(HRTF)在沉浸式音频渲染中具有基础性应用,但其高度依赖个体差异,随耳部、头部和躯干形状变化显著,因此需个性化处理以实现精准双耳渲染。近期,去噪扩散概率模型(DDPM)作为一种生成学习技术,已被应用于多种信号处理任务。本文首次提出一种基于人体测量数据条件的DDPM方法,用于生成个性化头相关脉冲响应(HRIR,即HRTF的时域表示)。实验结果表明,该方法在性能上达到当前先进水平,验证了扩散模型在HRTF个性化中的可行性。

原文摘要 · Abstract (English)

Head-Related Transfer Functions (HRTFs) have fundamental applications for realistic rendering in immersive audio scenarios. However, they are strongly subject-dependent as they vary considerably depending on the shape of the ears, head and torso. Thus, personalization procedures are required for accurate binaural rendering. Recently, Denoising Diffusion Probabilistic Models (DDPMs), a class of generative learning techniques, have been applied to solve a variety of signal processing-related problems. In this paper, we propose a first approach for using DDPM conditioned on anthropometric measurements to generate personalized Head-Related Impulse Response (HRIR), the time-domain representation of HRTF. The results show the feasibility of DDPMs for HRTF personalization obtaining performance in line with state-of-the-art models.

音频生成扩散模型个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。