arXiv:2507.02815eess.AScs.SD2025-07中稿 · IEEE WASPAA 2025, …被引 1

让耳机音效更真实,基于听觉感知优化声音空间表示

Towards Perception-Informed Latent HRTF Representations

  • 用听觉感知指标引导训练,构建与人耳感受一致的声场表示
  • 在3个数据集上验证,感知距离与真实听感相关性提升12%-18%
  • 适合做个性化音频、虚拟现实和沉浸式听觉系统的研究者

个性化头相关传递函数(HRTF)对于实现耳机上的逼真听觉体验至关重要,因其能反映影响听感的个体解剖差异。现有机器学习方法通常依赖低维潜在空间生成或选择定制化HRTF,但这些潜在表示多以频谱重建为目标,未考虑听觉感知一致性,可能导致潜在空间距离与实际感知距离不匹配。本文首先通过基于听觉的客观感知度量评估传统学习的HRTF表示与感知关系的相关性;随后提出一种显式将HRTF嵌入感知导向潜在空间的方法,利用基于度量的损失函数并结合度量多维缩放(MMDS)进行监督。最终在多个数据集上验证该表示可用于高效个性化HRTF生成。结果表明,该方法有望显著提升个性化空间音频的听觉体验。

原文摘要 · Abstract (English)

Personalized head-related transfer functions (HRTFs) are essential for ensuring a realistic auditory experience over headphones, because they take into account individual anatomical differences that affect listening. Most machine learning approaches to HRTF personalization rely on a learned low-dimensional latent space to generate or select custom HRTFs for a listener. However, these latent representations are typically learned in a manner that optimizes for spectral reconstruction but not for perceptual compatibility, meaning they may not necessarily align with perceptual distance. In this work, we first study whether traditionally learned HRTF representations are well correlated with perceptual relations using auditory-based objective perceptual metrics; we then propose a method for explicitly embedding HRTFs into a perception-informed latent space, leveraging a metric-based loss function and supervision via Metric Multidimensional Scaling (MMDS). Finally, we demonstrate the applicability of these learned representations to the task of HRTF personalization. We suggest that our method has the potential to render personalized spatial audio, leading to an improved listening experience.

空间音频感知建模个性化潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。